Pith. sign in

Paper Citation Record · LEDGER

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

As of 12 August 2026, this Paper Citation Record lists 100 of 125 outbound references and 59 inbound Pith citation observations for arXiv:2505.16933.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.16933 v2

Coverage vector

measured 100 of 125 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T03:46:06.074416Z

measured 159 of 159 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00

measured 59 of 59 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T16:51:18.515307Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-07-11T03:07:52.266883Z

Reference resolution

100 of 125 outbound references displayed

  • verified exact57
  • verified fuzzy40
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch3

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation eb3827e3-a287-405e-a693-ca44886b26e5 · outbound

This paper cites Visual instruction tuning.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Visual instruction tuning

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.504560Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:945a0b8aec3fc1cd5327ecddfe6972f92e63b0a192dce1e6f512ddaa617b2d57

Observation 1da84e56-3bd4-4ebe-ab82-4a7ef24e5818 · outbound

This paper cites Improved baselines with visual instruction tuning.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Improved baselines with visual instruction tuning

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.427693Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:5ae47154a533d94d5274e5901b7bcf0e5cad8c2fd1eda34e66831900ba746e46

Observation 3e14fcb8-15a1-4822-85bd-75e105605ac1 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning LLaVA-OneVision: Easy Visual Task Transfer

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.374982Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:04da2d9c7ca44e5f40225d0c4a2533ec64a9ce95783f080282a978ca8b307799

Observation f169e065-3353-4bc7-8af7-153c59d0da1b · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.432745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:7cddddf0e023fbc8ad3fd8423c13685447eab87e9b8c9dbcdc7cd124bcbf871c

Observation d0c19dc5-0c8a-4ca8-9786-8f45aa59691a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.378866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:7b9cb2e8d70379f1395eecef2471e53212440ce4e421b4cd0e7849423d536f11

Observation 4d8324a0-973b-47e8-b9cd-e6b79cde81f7 · outbound

This paper cites Chameleon: Mixed-Modal Early-Fusion Foundation Models.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Chameleon: Mixed-Modal Early-Fusion Foundation Models

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.382760Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:72b30d5432ae6a1018b0b3f3ef497a1c44f4840004d315c26342bc777aeb33f8

Observation 4ce34fd3-d640-4d04-896e-3421ccbaa926 · outbound

This paper cites Kimi-Audio Technical Report.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Kimi-Audio Technical Report

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.386499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:cf15f074946dfd2b898492c237ecf413ca4bc15785bd7d5421b4cb4da02fec90

Observation 8b25ced2-3437-4776-af5b-64709d4dc2ab · outbound

This paper cites Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.391139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:1cc090b240e8b0ef12b23dcce9524b1b44d086c35f9a28ddccdb0290ef07f971

Observation 06d1d233-677c-4da6-a2ad-61236a8de789 · outbound

This paper cites GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.395741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:3fef743cb3aed2fe9e3ad13a056fc7c4fc3a5bf3d049e512142cefe3cc0b0e52

Observation 40cb993f-0961-4de2-82ad-085c8318420d · outbound

This paper cites InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling

Reference 10

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.399421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:875b8b475e9b74de51ea2108e4c987439ad6ae3f3a8faca089499402886d31ad

Observation 166dcc9d-cf00-4591-be91-8ac64bfe4c22 · outbound

This paper cites Sharegpt4video: Improving video understanding and generation with better captions.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Sharegpt4video: Improving video understanding and generation with better captions

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.451055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:0430c80eff44c1ba23972c4fe29df745effa9938891f60ef014bf31314bf055f

Observation 331ed5d7-2084-415c-bfeb-bc831b564724 · outbound

This paper cites LLaVA-Video: Video Instruction Tuning With Synthetic Data.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning LLaVA-Video: Video Instruction Tuning With Synthetic Data

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.403500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:44861047a20e3094af78fd1d99e3446f014a2a987e80b496a53c4b9dd7f1603a

Observation b7a15082-b335-4343-b195-3bde33bcc2d9 · outbound

This paper cites Improving language understand- ing by generative pre-training.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Improving language understand- ing by generative pre-training

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.456509Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:0cad1cee0cb7e7f28678a6ddde00f0f9dfdc4414aa9e47bdd4433c827cc30b96

Observation bcba5c8f-c61c-4fd9-a7d7-35d5f3be3d9c · outbound

This paper cites Language models are unsupervised multitask learners.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Language models are unsupervised multitask learners

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.459411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:a3d9b50b21614ffb78aff138aca3ca9ebcebfc98c289db5cbac6fc72ccf6cc7c

Observation 48d415b4-551d-40c4-aef8-c91e1e3283ad · outbound

This paper cites Language models are few-shot learners.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Language models are few-shot learners

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.462167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:02e6c9a1664ba63495d7316da348c1b489a2720140bf4169638d3b948ac763df

Observation d9e3779f-853a-401e-92bc-4b97125945e0 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning LLaMA: Open and Efficient Foundation Language Models

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.407078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:fb548c71ae6bb726e766373154fa87343aef1b675f5d44fc3992e74dd90b7baf

Observation 9f4fe500-7a64-4a75-a2a4-354d104ff9bd · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.410445Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:63592ea61f71730992de729232f7c423b30547760153a8d5f40d87a0e28ac3c1

Observation 863a75d8-8b69-4ae0-8d25-7ca290de2a65 · outbound

This paper cites The Llama 3 Herd of Models.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning The Llama 3 Herd of Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.414170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:f714ea38e718696a811d28dcfa5681afedbefe40488cbe9abb91d9bb00f58a0c

Observation 9ea18293-1c8d-4055-b5df-92e8a040a93b · outbound

This paper cites Qwen2.5 Technical Report.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Qwen2.5 Technical Report

Reference 19

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.417644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:e94cd9d7085718662538b675e8ed87a6b467d4debf57a4ad0ad45b24c5a4c4c5

Observation 9799e762-59f7-4ad2-9188-2c2e8e84faa2 · outbound

This paper cites Textbooks Are All You Need II: phi-1.5 technical report.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Textbooks Are All You Need II: phi-1.5 technical report

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.421714Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:22d72e8123ca991a35f6ca3c9843e8ccfece1366dd66573fb56a5de6ae90cf21

Observation e20982b7-90c8-46b3-8f8a-598f0a197b53 · outbound

This paper cites DeepSeek LLM: Scaling Open-Source Language Models with Longtermism.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning DeepSeek LLM: Scaling Open-Source Language Models with Longtermism

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.425044Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:f9061acb7c04c6af474f1ae4dd681c7176f334c33df85052c9be2e970bf7c2cf

Observation 9d44a0bf-8555-41d1-9c94-7eee55cd6dfe · outbound

This paper cites Deep unsupervised learning using nonequilibrium thermodynamics.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Deep unsupervised learning using nonequilibrium thermodynamics

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.481351Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:5cfc934fc6e1a34cddd61819e48f8171d741518e2d86437ae03d427e82bea7cb

Observation e804b2df-89ef-4240-88a4-325bfb7d01da · outbound

This paper cites Denoising diffusion probabilistic models.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Denoising diffusion probabilistic models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.484167Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:d344c76a30b0d765ea29dc4ca96e8c4f05c1a97ec98458fba8fa6001fb4cb2ed

Observation 34e3f717-f466-4afe-9689-add636dc0769 · outbound

This paper cites Score-Based Generative Modeling through Stochastic Differential Equations.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Score-Based Generative Modeling through Stochastic Differential Equations

Reference 24

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.133700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:2b5ca5c38b9c300feef11ad6aefc6ca20fbac4f51fedd66b3a1cd07be2c3ab02

Observation 8edaaf77-deae-4cfd-8618-38e6b417e85a · outbound

This paper cites Argmax flows and multinomial diffusion: Learning categorical distributions.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Argmax flows and multinomial diffusion: Learning categorical distributions

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.489056Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:e525d4f4be8847934366133b5135a81c8480dfe98ac8385f166e1580eb6ea6c1

Observation 7c97355f-9f1e-43b8-ac90-3c6554c5ba9f · outbound

This paper cites Structured denoising diffusion models in discrete state-spaces.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Structured denoising diffusion models in discrete state-spaces

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.491672Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:a7d4f86ea49e9136af59dca56e313892f06e9e0497ab21a2124c55f3f9b3e4dc

Observation f52b048d-4123-4f13-83c6-bedd5d61d168 · outbound

This paper cites One transformer fits all distributions in multi-modal diffusion at scale.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning One transformer fits all distributions in multi-modal diffusion at scale

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.494360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:1b1607ebc699105f7f13dd6c14ff92e5e653ab18fa7336e0ce97e53bb46570dc

Observation 0223dedc-50b9-4d17-ab40-db591575af9e · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 28

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.139700Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:7842d33b727483e0670cc2ecc0feee920c61cdfd00a818143e169e143ba4a686

Observation 296627a3-06a6-4c75-89c3-b84188b0c60d · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 29

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.145032Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:3ba6937a171cf04f93eb00cf45e92b220c9a5ce303f33feb4e0186b09e607d83

Observation f69b31da-aeac-4421-b8fc-b816f8812f13 · outbound

This paper cites JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation

Reference 30

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:46:06.149969Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:c198c8b42910bbc38cdef71151f226fe83e78366bae2ec872fb24779d73eff81

Observation 839ae744-19dc-4705-a683-de022f9a29e4 · outbound

This paper cites MetaMorph: Multimodal Understanding and Generation via Instruction Tuning.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning MetaMorph: Multimodal Understanding and Generation via Instruction Tuning

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-17T07:51:13.882527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:97c87aaf23c91e5e646ac5b697f270ea42bcc850cbc4af083b4a0ad73279e831

Observation d0d52794-3986-4ec9-92e2-92310a2bf8d1 · outbound

This paper cites Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.155454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:b2bb0ea61e1326744446a262eb036c352bb64862394e6620d86511f180a6a536

Observation 3f4cc231-4116-41ba-a654-3b56db6a9fd6 · outbound

This paper cites Unified Multimodal Discrete Diffusion.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Unified Multimodal Discrete Diffusion

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.160898Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:9be0527473d65b69b1b36a1f24b94e7553827c67b6bd4ecdd1de9cc2c7cac30a

Observation b40dcffc-9bae-4483-bd3b-485d13ea5e07 · outbound

This paper cites Dual Diffusion for Unified Image Generation and Understanding.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Dual Diffusion for Unified Image Generation and Understanding

Reference 34

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:46:06.166172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:c8e260ac56f67284e43336393e29ce6fe3e67a38a0bbe3af46c2219c4ccf2d7b

Observation da41ae7d-9522-4068-a4ad-eb02fdbf6299 · outbound

This paper cites A continuous time framework for discrete denoising models.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning A continuous time framework for discrete denoising models

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.515783Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:8b8e405e857e153207dc59c283fc3ad4c2ec64629caad9d74ba22e5554f5556d

Observation d5937ee9-e1f8-45ff-aa30-2c14e0817adb · outbound

This paper cites DiffusionBERT: Improving Generative Masked Language Models with Diffusion Models.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning DiffusionBERT: Improving Generative Masked Language Models with Diffusion Models

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.170976Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:e59829ee00ac5500d9293b21ba10c548c5e8d4349314f1ada775628d1adaa10a

Observation 5568ff17-9370-4799-ab26-0f87be51fb50 · outbound

This paper cites Score-based continuous-time discrete diffusion models.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Score-based continuous-time discrete diffusion models

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.521210Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:1330b67aedf95c3cfeb7cc240350d4ec958e906c197c503e24a836927c8fc169

Observation d00570c7-a48a-45eb-9999-4f9bb61511b4 · outbound

This paper cites Discrete diffusion modeling by estimating the ratios of the data distribution.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Discrete diffusion modeling by estimating the ratios of the data distribution

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.524078Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:c688ad493b45c9858a9b982d369e4fc64f8b092c6669d651310e542dcb7f501c

Observation c524b596-ef42-44fe-9b85-ed2d972577e3 · outbound

This paper cites Simplified and Generalized Masked Diffusion for Discrete Data.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Simplified and Generalized Masked Diffusion for Discrete Data

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.176281Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:6124e8ec93b83f78bf8a5f5e9400d58d6f824876da9fd0aaa8923b4e4ed5c462

Observation e7475394-6ea5-4e0c-9ed7-246eb2dc315c · outbound

This paper cites Simple and Effective Masked Diffusion Language Models.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Simple and Effective Masked Diffusion Language Models

Reference 40

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.181114Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:7227ca45983326b4391c51e45614d8b597406ea964d5f1aa098e156ea6e1ded1

Observation 6acf96e2-3ce4-45cc-9a1e-10feeba17223 · outbound

This paper cites Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-05-18T12:06:09.375609Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:19d0c14b276b3d80b55369f6ebbc7aeeeeaa2777ed299452544e033037970fdd

Observation 9bd7c7d8-9e2d-4f4a-8089-1341606f43fb · outbound

This paper cites Large Language Diffusion Models.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Large Language Diffusion Models

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.190689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:8e602e0ba2184bd20e00d0820d8057ec26999e0fddde150b78f02d4d3d1da9f1

Observation bd0b8896-2c9f-4690-baac-3ec50bfdcd6c · outbound

This paper cites Effective and efficient masked image generation models.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Effective and efficient masked image generation models

Reference 43

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.196182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:f288d2f923ef230c43fed967ad2ec5488db775bccfcf13f72cb4edc02b904577

Observation 0c0d6acd-f0e8-4f5c-b306-bf8ba73ece36 · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.201287Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:3f4f7253762d58b3fcead513065e5ef4e5484ef7505be672b7c4af49b50eb5fe

Observation cc9b7f28-f032-498c-b772-ba0ac2edc3dc · outbound

This paper cites Are We on the Right Way for Evaluating Large Vision-Language Models?.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Are We on the Right Way for Evaluating Large Vision-Language Models?

Reference 45

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.205879Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:b7005791cd3a7ffa55f4482f088a0a5a97e2ed12be815313f5c81494d0f849f8

Observation 445b6210-ae1a-4947-bae5-b525d1df4109 · outbound

This paper cites Scaling up Masked Diffusion Models on Text.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Scaling up Masked Diffusion Models on Text

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:46:06.210623Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:fe7db4d07081c8825f90346538b678e112127a1c2071a00e4923e9769efaf68a

Observation 28decc83-d2eb-411a-b305-0dbd6b0f38be · outbound

This paper cites Generative flows on discrete state-spaces: Enabling multimodal flows with applications to protein co-design.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Generative flows on discrete state-spaces: Enabling multimodal flows with applications to protein co-design

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.551165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:a2399cb114b3228998e7ded6ce46155427bf4d082057ef53e4c83293897e3d14

Observation 08656729-0c85-4f49-a871-8b13c4e11075 · outbound

This paper cites [MASK] is All You Need.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning [MASK] is All You Need

Reference 48

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.215090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:800c14a661113b60ed2488b13048041e85a9fd6e0090888e8b9d84e1e89278a5

Observation 41c3e79d-43c6-47d2-a54b-8e3497ef3a44 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Learning transferable visual models from natural language supervi- sion

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.558145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:8891c14ce65f1f06d692941729e14cabfee31107eca82e1ff175e75e76b72975

Observation 235af7ba-d005-4b4e-a10b-404e509d1973 · outbound

This paper cites Sigmoid loss for language image pre- training.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Sigmoid loss for language image pre- training

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.561415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:e70b8f16f997da21ecbfcae1c2534ece68f85f758fb551430c69f160d55875b0

Observation f3554d0f-20eb-43be-b3ad-ef62fd9bae44 · outbound

This paper cites Wan: Open and Advanced Large-Scale Video Generative Models.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Wan: Open and Advanced Large-Scale Video Generative Models

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.219339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:a4f162e479a7390e348df1f5b4f5d2308f8cfe455fdae84bf2e71fffa2b6197c

Observation b7916aac-0447-4d56-aceb-d18a2a28da93 · outbound

This paper cites CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.223481Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:f3484a03d62d20b5b2d4126f899662894d7accf05fd6b6b543ceb0fc92c149c1

Observation 43fc7807-9237-4cae-b9f5-66485888a76e · outbound

This paper cites HunyuanVideo: A Systematic Framework For Large Video Generative Models.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning HunyuanVideo: A Systematic Framework For Large Video Generative Models

Reference 53

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.227621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:fc52c80a05e1c57b76e3f67734fdbce88b38c07cb7e105929480e8631c231bcc

Observation 97d3e00c-f044-47ea-95a8-d494f548746f · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Llava-next: Improved reasoning, ocr, and world knowledge

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.572843Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:276164c1cb63c766a16919678c2e5b84e84a1213ed965845e49357f952006e38

Observation 74420e64-e9bf-490e-9bad-8518905b5848 · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 55

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.232126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:addfa2d6fb71f9202e5eb58dcbbe40c303b481f3addc8ede6dd4e6c34ec979ec

Observation 56107efc-5ba8-45f9-8914-68e61b75a389 · outbound

This paper cites VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.236748Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:4f599252643019c8f7cc63cc9338b2b1b9083f76548af02f64ed0225f4a832b9

Observation 766fedce-a1e5-4383-85b2-bef96fdbbfe2 · outbound

This paper cites Qwen3: Think deeper, act faster.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Qwen3: Think deeper, act faster

Reference 57

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.430142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:78e26428e6c5a631e0602273f4fc88f0ac251f8b48115a306f84db36dad85199

Observation b09d0fa8-ad5f-4e5d-b554-7ded94b6979a · outbound

This paper cites Proximal Policy Optimization Algorithms.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Proximal Policy Optimization Algorithms

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.241029Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:605e827ac9bbb1f1559f6bd3c380d4133bb9fe1fe9ea65555091cc59b04f835a

Observation e69c9a60-9b6b-46eb-a1da-eaf0754f3d8d · outbound

This paper cites Direct pref- erence optimization: Your language model is secretly a reward model.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Direct pref- erence optimization: Your language model is secretly a reward model

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.437948Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:334b506ad4afd1f655a0206b5680a7e589c85d6ffc35144e63edd25ae67c09c2

Observation ae8df455-53cf-436d-b6a1-2d4c87990ae6 · outbound

This paper cites Simpo: Simple preference optimization with a reference-free reward.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Simpo: Simple preference optimization with a reference-free reward

Reference 60

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.440182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:d94b52ac6a739d9cada2fb3f5c44d92fbd49c5670fda97f9970070d38e6e7fcc

Observation 25e23fe3-dd3c-4c96-8eb7-124c01a569c0 · outbound

This paper cites DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models

Reference 61

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.244804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:4f1adc12f1d27522f47f5846bf3200f983d51af64846d29578a51b3ea6d11163

Observation d357fa1e-e6c9-4c78-9844-a96a9eac384f · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 62

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.445530Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:ea8aeaea934d6e7ce72d3c32ef30de2c5b7b8443a8072dccfc13bf0737220a5f

Observation 59aa0ef3-eed0-4f7e-870d-2f320e92a188 · outbound

This paper cites MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark

Reference 63

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.249031Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:177f55784802c51a30c25cd4609d38d25492d9e369349d5d0d24dbd11160cf1d

Observation 954aae21-7be7-4b15-811b-87e85925fe7b · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 64

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.253398Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:356e6d7f30ba5b7a117491b958635678d3d622b9beb4f14c9f4b82c1f9920cd4

Observation 96b9674b-7c79-4d0c-8582-4f29ae20bde2 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 65

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.257804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:d53754f1efea931b025ba0c496e4ce28ba9e3483fae088893f3c3333927ebb37

Observation 94533f71-da6b-4817-9cd0-cfd613a27ac5 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player?.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Mmbench: Is your multi-modal model an all-around player?

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.467768Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:c6664b641494d658ac012c9e79dfc535b0c3c64d9d7dc2888eb44b956c65abf3

Observation 6d754cfa-d9d4-4afd-adff-19a806badf15 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems?.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems?

Reference 67

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.470591Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:b630028d566a1999b288d79adb9c787bc7b0798ff8af4ce65b48acf4c0bf0a31

Observation 0dfe3396-9f15-45bb-a736-7942eb673620 · outbound

This paper cites Mathvista: Evaluating math reasoning in visual contexts with gpt-4v, bard, and other large multimodal models.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Mathvista: Evaluating math reasoning in visual contexts with gpt-4v, bard, and other large multimodal models

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.473458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:3d3e5659afb3cc0acf01662bef7abbc41e7e713e305741e39089a26196936f04

Observation 9dd7dea5-c8ce-406a-a6f2-67c4e0c9f396 · outbound

This paper cites A diagram is worth a dozen images.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning A diagram is worth a dozen images

Reference 69

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.476384Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:de576aeebbc53888ff93d27f1140e924022fcd256c631c19a768d3d78074b55e

Observation f6c0efea-4f43-47c8-a549-b2bac16b1c8a · outbound

This paper cites ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning

Reference 70

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.262206Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:94269ff624ec089c35d922f54691d32bde6613b04aeec4f5f9e74c3aec193253

Observation a32c1cb8-8f25-42d2-817d-2bf52fdc48e3 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Docvqa: A dataset for vqa on document images

Reference 71

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.486618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:55c3bda51dd3e3a8fd0b90e8988b6080a3bdf357cf50759aba4c8ea26f1e432b

Observation 097ac91d-3457-4070-9e4e-facd08cb09ef · outbound

This paper cites Infographicvqa.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Infographicvqa

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.496875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:ce5077b8209aac8947db17a94f767d13a34cb82f45e7e9b675b8fd910402b0fd

Observation d4403c52-757d-4631-9544-b237b722ab18 · outbound

This paper cites Grok-1.5 vision preview.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Grok-1.5 vision preview

Reference 73

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.499231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:3b10f11e5ab9f7abad6405c2590ea8445231bac5713b30e9f63213d7c8b025d4

Observation 546a3dee-5b6d-4675-a23e-1e94049a767d · outbound

This paper cites MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding

Reference 74

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.266916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:38e7f75a3aafdccdd2d1f24baab7f266495169c4b597380b7293355f15013fce

Observation 34eaec19-4036-40b3-813d-b147c32d8412 · outbound

This paper cites MLVU: Benchmarking Multi-task Long Video Understanding.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning MLVU: Benchmarking Multi-task Long Video Understanding

Reference 75

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.271119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:815c2a98c3db711869ac3e3c7a00363defcec595f7d4134aa8b27bb44aaf3f9c

Observation 4362463a-c272-4aa9-9fa4-c71980ac5c6c · outbound

This paper cites Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis

Reference 76

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.275850Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:1054de128c3dcd745fd343a44204bb974d4b86070b48004f54421d8bae67c8b9

Observation 079da712-36c7-4575-8ad6-f9906f60615e · outbound

This paper cites Sharegpt4v: Improving large multi-modal models with better captions.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Sharegpt4v: Improving large multi-modal models with better captions

Reference 77

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.512512Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:5aea896ef1d9aa9f1dede5739fc6624b97c7e7b793827f99bd802b01778cf560

Observation 98b271ca-217a-457f-99b8-2aae51ed6882 · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 78

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.518834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:fbae482d5d2a317eb5c1534c6c051ea7449ac37dfc514a21ad986091a8be2436

Observation f926c3c7-f79c-4371-a099-c6908596a461 · outbound

This paper cites DeepSeek-VL: Towards Real-World Vision-Language Understanding.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning DeepSeek-VL: Towards Real-World Vision-Language Understanding

Reference 79

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.280227Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:47227d69dca4f6f49775f9451b6b8bf54ed1b595c880ac868a49ce2a71b2a519

Observation 49a6cef4-ddf7-432d-80ca-9b727dc6395c · outbound

This paper cites DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding

Reference 80

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.283985Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:22ddf6e15e63cc5e5309222a415d343ad9b44d7e627e0b26d8887c166baab987

Observation 010a517e-f0b1-4adc-9d7c-92ee868653fc · outbound

This paper cites Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation

Reference 81

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.289016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:ccf942f94808a626f7966e8382fb6b471ef8f161acd6630c609708ecbbb6e876

Observation 415cfe4b-de35-47ff-9419-73ac225adc98 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.292960Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:aea62467950345024c1f9baeb0b633ced91a0f43461b4e467091ae0c16851298

Observation d483a297-5ab9-499c-b3ed-e7815f5a501e · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Emu3: Next-Token Prediction is All You Need

Reference 83

Resolution
verified exact
local_arxiv, observed 2026-05-17T03:46:06.296588Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:ee4724e69bcd7171206b23a0db98141ebcbcdfed70ef9a2c7765d67e84748f99

Observation 90c34c26-3cb2-43a8-9c6f-4cb44e5ec9cd · outbound

This paper cites What matters when building vision- language models?.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning What matters when building vision- language models?

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.541997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:72bcbd7e21cba0e2dc15073ab0347a438d8711a18d95a1a6b769ec87c0f6248f

Observation 6335c45f-b843-448d-85f7-16fa449c88e4 · outbound

This paper cites Diffusion-lm improves controllable text generation.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Diffusion-lm improves controllable text generation

Reference 85

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.545358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:a650151ff6dac5fa4aad2e947eb2aa80f0e31e121e8e9f238ff37095faf1350b

Observation 06e83cfe-da1b-4941-b442-33935c749004 · outbound

This paper cites DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models

Reference 86

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:50:38.345995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:b67f42aad25afe3b46173b24fc576058cd4eb6895ccce2f2ec29dc0112b2c927

Observation b375bf3c-173b-4363-ab3b-cb69fb510164 · outbound

This paper cites SSD-LM: Semi-autoregressive Simplex-based Diffusion Language Model for Text Generation and Modular Control.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning SSD-LM: Semi-autoregressive Simplex-based Diffusion Language Model for Text Generation and Modular Control

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.304802Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:9819fb9d44dfeccf5e4f8702e8021e0575865aa7838c52cb658fdeca23a9f202

Observation 5e5f35e3-e498-4783-88d0-c5f65f595e2b · outbound

This paper cites Self-conditioned Embedding Diffusion for Text Generation.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Self-conditioned Embedding Diffusion for Text Generation

Reference 88

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.308600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:f2806716ec3b26a338cfbbc6eba485a205106a07680cb7cb520b25a27379f6d5

Observation 826ada18-bbca-486e-abb0-90a06845a8e1 · outbound

This paper cites Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning

Reference 89

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.312461Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:7eff1458ed630726e0255142a8715d0918efc2d525534b17b66977bb5ac0fb44

Observation aee58e8e-e87d-46ec-b43d-2163b6a6cd53 · outbound

This paper cites Continuous diffusion for categorical data.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Continuous diffusion for categorical data

Reference 90

Resolution
verified exact
arxiv_id, observed 2026-05-18T03:30:22.553521Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:7a8b154f8dc7e200d6c15eb48d81071a501dc92c289f5ef097166a1c1b767215

Observation 335ae4b8-e1ab-445d-8b78-1d88a7ac3d12 · outbound

This paper cites Categorical sdes with simplex diffusion.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Categorical sdes with simplex diffusion

Reference 91

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.575634Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:9cf23abff2c77e0a76d41df36e24a608f22b4db63a627aa28921086de81f21d0

Observation fc0cb21d-4572-43b1-9348-b787ecbec784 · outbound

This paper cites Ar-diffusion: Auto-regressive diffusion model for text generation.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Ar-diffusion: Auto-regressive diffusion model for text generation

Reference 92

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.578941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:94105b1f173df43b4df03acbac366f0f1ff28265a68df2b3e01ad3ee0d69bc1d

Observation 79705010-b49f-4211-9800-40a077a77b6b · outbound

This paper cites Tess: Text-to-text self-conditioned simplex diffusion.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Tess: Text-to-text self-conditioned simplex diffusion

Reference 93

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.435422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:16d16934d64530cf89a1f5e0ea2361f08e47701584d949c54274717b582627ab

Observation 9f78f635-7cfe-4958-9aa9-9718fecfd85a · outbound

This paper cites DINOISER: Diffused Conditional Sequence Learning by Manipulating Noises.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning DINOISER: Diffused Conditional Sequence Learning by Manipulating Noises

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.319703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:719f31bfd42f21528168551b925bc5b6ab022922a1f357165697e854f3977ed6

Observation 89cc55f2-661b-454f-b7c3-701027ffb29c · outbound

This paper cites Planner: Generating diversified paragraph via latent language diffusion model.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Planner: Generating diversified paragraph via latent language diffusion model

Reference 95

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.448323Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:24c784c8f14ba719975c5d7f07e79d41699a7277c6cc5969416f701f71b1e578

Observation cd448a65-09ea-4803-9e1f-529f3b07ec74 · outbound

This paper cites Reflected diffusion models.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Reflected diffusion models

Reference 96

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.453674Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:638ef368f8658da7418f85445f0eb750c336114fbf2f954a306acf181651e120

Observation b78038ef-7bf1-429d-b8f5-e0cc015bc126 · outbound

This paper cites Bayesian Flow Networks.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Bayesian Flow Networks

Reference 97

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.323058Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:019a462fe939a660c4fd2c9901a36a48423913391a485697a3db99ae17be69df

Observation 2f7b1ba1-11c5-47ea-ba0a-1e63ff2b25f8 · outbound

This paper cites Text generation with diffusion language models: A pre-training approach with continuous paragraph denoise.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Text generation with diffusion language models: A pre-training approach with continuous paragraph denoise

Reference 98

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T03:46:06.478931Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:f5c3799acd276efd20df58dc80db3f189541499302ea3e3d6e952c1e989ad4af

Observation 82d4db9d-1bb7-442d-a71c-07dc8d9c0f1e · outbound

This paper cites Unifying Bayesian Flow Networks and Diffusion Models through Stochastic Differential Equations.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Unifying Bayesian Flow Networks and Diffusion Models through Stochastic Differential Equations

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.327035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:fe7eb037e6a4d1314288f79aa25d15b21e45aa01d031c2f4efff678891b228b5

Observation c40fe4d7-3a34-4b71-927d-6fb9afbb34c5 · outbound

This paper cites Target Concrete Score Matching: A Holistic Framework for Discrete Diffusion.

LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Target Concrete Score Matching: A Holistic Framework for Discrete Diffusion

Reference 100

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.330844Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-17T03:46:06.074416Z digest=sha256:74a673d7b9acbe6b811ccb9b8febb0e6a31ad7fe9c2efbc48389f624483c54d5

Pith citing papers

Observation c7785dfa-a5d3-4648-ae12-9fccc85b42cb · inbound

Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding cites this paper.

Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 38

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.580274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T04:28:02.231373Z digest=sha256:e9a8feebaa7e098923723c410d211d38c87f65caca0ed1f3e3de3822d8075001

Observation a82829d2-a820-4e1e-a799-5096c180d32f · inbound

The Philosophy and Physics of Duality cites this paper.

The Philosophy and Physics of Duality LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-06T05:32:07.588327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T05:32:07.588327Z digest=sha256:4a2a416fea76849def1c36e699bb4df5ccc0e800877eb02066856fe3d92dcd0d

Observation 07b2dcdb-1cff-428a-9a8c-6e1035ba0070 · inbound

A Survey on Diffusion Language Models cites this paper.

A Survey on Diffusion Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-05T20:15:15.869197Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T20:15:15.869197Z digest=sha256:79aa9675c24963b56eeba42192e47ae7c7aeec2ef191677856146a64f045dff3

Observation 072f861c-8656-4bc4-935d-2ad3398040cc · inbound

LLaDA-VLA: Vision Language Diffusion Action Models cites this paper.

LLaDA-VLA: Vision Language Diffusion Action Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-04T22:55:30.063762Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T22:55:30.063762Z digest=sha256:03b8a1a846adfeb9c2ae798e40816b6515c85aee1bb577ac9ea34de257cdffa3

Observation dca79717-210d-479b-90c9-37e24c8d55fc · inbound

Inpainting-Guided Policy Optimization for Diffusion Large Language Models cites this paper.

Inpainting-Guided Policy Optimization for Diffusion Large Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-04T17:57:47.054983Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T17:57:47.054983Z digest=sha256:0a44050594c0f8962117eae63ce205d9b101afd9b46699a90f9cd049f15ce50f

Observation 24b21a0c-3418-40c6-8124-7baa99028163 · inbound

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation cites this paper.

Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-04T15:39:42.150374Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-04T15:39:42.150374Z digest=sha256:1eb1febab4f159aee8534278622f4ad1eed87a3e521de9e6bf4dbc87ae38017e

Observation 40f9660a-6be0-4cce-93c0-e954c9b8ac14 · inbound

CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Credit cites this paper.

CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Credit LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-18T09:11:10.026064Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-18T09:06:25.049493Z digest=sha256:53e2e82ebf229e028a2d1768c9c66d3aa25561016bf9193128dd9a4de5e4cb2d

Observation 02d7388a-13b4-4c4f-b1fd-67961e2f98f9 · inbound

CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Credit cites this paper.

CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Credit LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-04T11:16:47.440982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T11:16:47.440982Z digest=sha256:95a4fde8cac4ae4d1f83282f73b48959ee50418cc61ee9bfded6721f34ea76a2

Observation 87f4633e-4417-412e-8e99-a5245e57c9db · inbound

A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models cites this paper.

A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T21:31:36.816979Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:31:36.816979Z digest=sha256:ac8b1ec8e282c48727b0b0c8674245a72a61561ca9a83323bee1f21f60145c64

Observation de50c696-4650-4c37-a92b-b230fba37457 · inbound

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models cites this paper.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:38.516268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:38.516268Z digest=sha256:aab521292ea232e5946c6bf3c8de8dfdb0a64cf960d20edb1508964aacd2c7d5

Observation cb444cf3-6c48-4bd9-ac63-0a6dd390f5df · inbound

Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in Speed cites this paper.

Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in Speed LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.580274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T22:29:08.669964Z digest=sha256:57796eba23cf3ef152d7500da18c2e5e01686b80a4aa194276a78adc413b5cde

Observation 7312622e-1578-4eea-aaf4-5a9bc24a6733 · inbound

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models cites this paper.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.580274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:827a1e1f76c7d8631ec3c2387681796e2d1125962e1e4dfdcadc8136db202fe4

Observation 91116f92-a9f5-4328-b953-eb8627c20fa7 · inbound

Streaming-dLLM: Accelerating Diffusion LLMs via Suffix Pruning and Dynamic Decoding cites this paper.

Streaming-dLLM: Accelerating Diffusion LLMs via Suffix Pruning and Dynamic Decoding LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T08:14:37.068486Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T08:14:37.068486Z digest=sha256:828759d7591e4dfa8a12f43a7c1e76fec4a2a15ab0695045b6fbe8416e05ece7

Observation 36b693b4-2464-4076-a5c0-d10994d5788e · inbound

DODO: Discrete OCR Diffusion Models cites this paper.

DODO: Discrete OCR Diffusion Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T22:27:44.206481Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T22:27:44.206481Z digest=sha256:4b5607489110c7040ff5e4b6e595d3740deddc4f64ebd3127643f5f3bcd297d2

Observation 1af74f9a-ebed-46be-bb48-014677dea550 · inbound

Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion cites this paper.

Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 49

Resolution
unresolved
no resolver link, observed 2026-07-15T13:43:40.241796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:43:40.241796Z digest=sha256:9e0373488a3b66de873381e5f61d921e85246ea727e650c6e8796f6dcddfd7d7

Observation 92a023c9-53c7-496f-a97c-b996afba0642 · inbound

Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models cites this paper.

Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.580274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T19:06:45.417233Z digest=sha256:c84e308b69e97bce6fe788b102ee6af75f95b1f1206dc5f4c990f0b55fb5a880

Observation d5b250fd-1231-4206-9f3b-fbbad600a124 · inbound

DMax: Aggressive Parallel Decoding for dLLMs cites this paper.

DMax: Aggressive Parallel Decoding for dLLMs LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 94

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.580274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T17:58:17.880199Z digest=sha256:8ea725c1d760435e073f1be965d5c2575dfdbb85d8ccfc9682be978d6a0f93b5

Observation 48172090-4760-4a73-96a6-c70627854ba7 · inbound

DMax: Aggressive Parallel Decoding for dLLMs cites this paper.

DMax: Aggressive Parallel Decoding for dLLMs LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 94

Resolution
verified exact
local_arxiv, observed 2026-05-19T16:47:40.114126Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T16:46:56.743268Z digest=sha256:7989c1f5695eae7b6a6e8fe56ccbd319c743e0df749d98fa0d5bda2bd8511ace

Observation 6ebb8afe-175f-4c15-93df-684a05f7cbbc · inbound

ECHO: Efficient Chest X-ray Report Generation with One-step Block Diffusion cites this paper.

ECHO: Efficient Chest X-ray Report Generation with One-step Block Diffusion LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.580274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T17:54:54.441128Z digest=sha256:89ddec737cc069072abcf76be996fbe1c7ad51ccc35884625ee5f2e402f6f486

Observation 401f5876-6c26-4a02-85cf-c91df8f2ce99 · inbound

ECHO: Efficient Chest X-ray Report Generation with One-step Block Diffusion cites this paper.

ECHO: Efficient Chest X-ray Report Generation with One-step Block Diffusion LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 57

Resolution
verified exact
local_arxiv, observed 2026-05-21T08:54:05.875779Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-21T08:52:04.531938Z digest=sha256:07c56a2f58c5e77e40900905549b1c4501bc75da4b2239b063a4026094763f56

Observation 52b44d06-1538-478c-8960-9592f83ddfaf · inbound

LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation cites this paper.

LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 56

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.580274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T16:08:58.648967Z digest=sha256:300e61ccf828b1fe55aead92c5500f6b3fedd10b55fa47f46471acb13fd05e8b

Observation 10beb142-c2e7-4b3f-b268-6d919da7e65c · inbound

LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation cites this paper.

LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-02T16:27:21.284010Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T16:27:21.284010Z digest=sha256:da5c53bf49684146ce37781cbc1ff2230e8a3dedc3203ed99fab10e4bc7ba2e5

Observation 1686a21a-cb41-41e2-9221-edc480f3f1df · inbound

Dataset-Level Metrics Attenuate Non-Determinism: A Fine-Grained Non-Determinism Evaluation in Diffusion Language Models cites this paper.

Dataset-Level Metrics Attenuate Non-Determinism: A Fine-Grained Non-Determinism Evaluation in Diffusion Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.580274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T13:31:46.940449Z digest=sha256:ae79f9a2f66f96199c3f1a63aa64c5d8f0649f7fc829237d8dd217c0de56f0c7

Observation c15f2d8e-0392-4f8f-84ee-744cf0776978 · inbound

DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis cites this paper.

DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T20:50:53.567415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T20:50:53.567415Z digest=sha256:76e3ed5b67bc90aa9d08ae88d525fda7b8e676c2eba745d0f755a4d2dfd8ad7f

Observation fbbafb58-abfd-4d16-b8ca-a2b898647d8e · inbound

BARD: Bridging AutoRegressive and Diffusion Vision-Language Models Via Highly Efficient Progressive Block Merging and Stage-Wise Distillation cites this paper.

BARD: Bridging AutoRegressive and Diffusion Vision-Language Models Via Highly Efficient Progressive Block Merging and Stage-Wise Distillation LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.580274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T14:28:47.505755Z digest=sha256:d8f84cd116c0d1e796f496a35d12a0deadf2d8f4c43a40cecc333276cdb61085

Observation eb4bde6f-c13e-46f7-a290-c7abca04ea42 · inbound

Stability-Weighted Decoding for Diffusion Language Models cites this paper.

Stability-Weighted Decoding for Diffusion Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 16

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:46:06.580274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-10T06:12:37.804816Z digest=sha256:dc0596c68574c94efdc20c71e7df53cb3c6c2c4e61074221789244de030159a7

Observation ca8d6825-d1c7-47a7-aaf5-00c0fd0daf9a · inbound

One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models cites this paper.

One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 173

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T03:46:06.580274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-10T04:56:35.796962Z digest=sha256:e410850f0e5f0fdf0c97f3a905a751cbf24baee057e0c57a33557ee9effa2b2e

Observation 0edd5724-900e-418f-9547-dc81ff4795ad · inbound

Continuous Latent Diffusion Language Model cites this paper.

Continuous Latent Diffusion Language Model LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 105

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.580274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-08T10:04:09.646578Z digest=sha256:d0865c7da317f7a76ba64515b740ad74c50908ba6da9a5ab67d7ae9ffc58483f

Observation b34a8d5b-8b14-4ff0-8664-9ddca8b66942 · inbound

GPO-V: Jailbreak Diffusion Vision Language Model by Global Probability Optimization cites this paper.

GPO-V: Jailbreak Diffusion Vision Language Model by Global Probability Optimization LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.580274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-11T01:48:09.652599Z digest=sha256:0dccc0f632bce7b79e8e47ed2dd277682b03956472a7f8a65d50cba6748b8941

Observation c57969a2-7928-4a43-bffd-f52edbfd67af · inbound

GPO-V: Jailbreak Diffusion Vision Language Model by Global Probability Optimization cites this paper.

GPO-V: Jailbreak Diffusion Vision Language Model by Global Probability Optimization LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.580274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T04:13:54.794576Z digest=sha256:0eb663e383980d0a99a5f7837e517aa1a15b7fecef2397025abfd8c9a734d70c

Observation 3a02f2a6-a770-466f-81a8-1cb2f278de4c · inbound

Relative Score Policy Optimization for Diffusion Language Models cites this paper.

Relative Score Policy Optimization for Diffusion Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 99

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.580274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-05-12T03:47:42.196931Z digest=sha256:18b463f1263e4fdecde784eefa1f89677c666c2d563affc892d20fd96d8c4920

Observation 9a8fe905-5117-403b-91d0-fdf1754ade04 · inbound

ELF: Embedded Language Flows cites this paper.

ELF: Embedded Language Flows LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.580274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-12T03:30:11.905667Z digest=sha256:a5eba24443d16a75e42fc9c56b592388e1ff8b3fc5f630396f9448faac90c13b

Observation f89f8dc6-cc6c-4f68-9a6f-f28f0e3c50d5 · inbound

ELF: Embedded Language Flows cites this paper.

ELF: Embedded Language Flows LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 78

Resolution
verified exact
local_arxiv, observed 2026-07-01T14:05:46.219054Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T22:21:54.815122Z digest=sha256:dca5596138744ca68c7912e214252a5bda45db323492f59c0cb80ef8d1c4b503

Observation 0541e5a3-d182-4862-882c-223458aa6288 · inbound

MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving cites this paper.

MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.580274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-14T20:54:50.887381Z digest=sha256:b0aac86045f72567314f33f071729e0fe7940b476baacf7b2219e9534f1c4547

Observation c046c1a7-c9b0-4ecd-b767-f7c846f2431f · inbound

MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving cites this paper.

MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.580274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T05:07:58.953866Z digest=sha256:d531b387f5f494f3630a78843e6466423194804d87112a310d7275172b3e994a

Observation 8a319274-e01d-4b42-b589-99229e3338c1 · inbound

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models cites this paper.

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.580274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-15T02:14:59.487486Z digest=sha256:761a87721632cc78eeecfe614cdb697e3fe38a876b60ba965332dfe72bfb4e0f

Observation d91d7967-7d2b-4495-9d9c-4fc7050ec5a9 · inbound

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models cites this paper.

Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-20T21:49:05.340916Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-20T21:47:32.112057Z digest=sha256:a5f16b13108c59fcbba637d9a6493e1da554ed5f9b9747089b040ac861198c47

Observation 97d16801-5838-4526-827a-d903d58cf0e7 · inbound

Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models cites this paper.

Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 44

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:22:48.437357Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-19T21:18:03.005508Z digest=sha256:7ed428262f9052f10c3b9aaf662f5e30ef7d989c86e3f2b46966219c9bd2246f

Observation bf27191e-27c8-476a-a3fc-64e88f6847c7 · inbound

Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving cites this paper.

Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-05-25T05:06:38.510406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-05-25T05:00:58.024124Z digest=sha256:c543b15d47cbc06f38fbdd34d9ff7371e5b09f77ba54b43c3d9e8fe277617d61

Observation 941648a2-a284-49f5-8c8a-8a06e1b960ee · inbound

Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving cites this paper.

Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 23

Resolution
verified exact
local_arxiv, observed 2026-06-30T16:35:12.292159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T16:34:21.422620Z digest=sha256:39fcd4ff10d895d7c77582381dfe031cb744410b7575ef8fdc5763cce3cf4955

Observation e3c9f15a-40bb-4901-940c-6be7a95ad9c1 · inbound

Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models cites this paper.

Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-06-29T23:14:01.649826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T23:09:32.594194Z digest=sha256:a9a7382134105a203e6933c18ba28634645367e78b2774ad0dfcbb133f20381c

Observation f822b833-a84e-436b-b1f1-f51754cfbdf7 · inbound

AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling cites this paper.

AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-06-29T08:53:15.643977Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-29T08:52:10.642943Z digest=sha256:8edaadf3452e1b2566bfba824ac8987acb58fb0aa1c0a9d1cc74002ff3f4fb7e

Observation 9ae228eb-72a0-4e85-a531-164d1de881bb · inbound

dMoE: dLLMs with Learnable Block Experts cites this paper.

dMoE: dLLMs with Learnable Block Experts LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 59

Resolution
verified exact
local_arxiv, observed 2026-06-28T22:52:45.042761Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-28T22:50:51.900169Z digest=sha256:98bf1b6e0041c8aab73144a299a3fa375f2a0e4db2efc6fa26ee2de8189eb7cd

Observation ff7ad414-ffcb-43a5-b5cd-3bb6c841210e · inbound

Dynamic Infilling Anchors for Format-Constrained Generation in Diffusion Large Language Models cites this paper.

Dynamic Infilling Anchors for Format-Constrained Generation in Diffusion Large Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 82

Resolution
verified exact
local_arxiv, observed 2026-07-02T08:06:48.409248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-28T06:20:27.041099Z digest=sha256:b17fb879ea52c84abdaf90aee04f30df17425e66064e886561b9b41f36786fb5

Observation 930ae589-b850-42be-8cc0-a8501c1e2886 · inbound

Prefilling-dLLM: Predictive Prefilling for Long-Context Inference in Diffusion Language Models cites this paper.

Prefilling-dLLM: Predictive Prefilling for Long-Context Inference in Diffusion Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 14

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T04:57:38.705417Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=arxiv_source observed=2026-06-27T13:30:19.620692Z digest=sha256:5fe6f17e50bb0b1e1d1606e66cabe9b6037009659198a9b3dd6e9ce4d1ccca0b

Observation d13cc417-d2c5-4d66-8cfd-cc2f299eca4d · inbound

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models cites this paper.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:18.047342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:539deaf9b1ce6ef51678b5b8d54d673e3572ac9f67b2464950798d7a975f27f8

Observation 32551ffa-a4de-414c-8aa5-6147c706dc9e · inbound

Improved Large Language Diffusion Models cites this paper.

Improved Large Language Diffusion Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-04T19:20:05.614967Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-25T21:34:45.620481Z digest=sha256:a98e4d1b36a48ee6d6dc4e6ab2fe3ad2d25dbd8b716ea77e25147ba0f6d44cc2

Observation 02c0a308-8bee-444f-8aa6-fc63c7094570 · inbound

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis cites this paper.

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 52

Resolution
verified exact
local_arxiv, observed 2026-06-30T08:14:26.650220Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-06-30T06:07:55.600338Z digest=sha256:c74e983b59e62a87b86066f127921e4121692f58cd4c3e1abee963397ffd5376

Observation 09b26b26-9119-418d-bf41-2f6e9082c521 · inbound

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis cites this paper.

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-02T09:39:38.012846Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:39:38.012846Z digest=sha256:8d57c284134f95987852558496e0235a15c7a4387326e1e4d01a2464badca8b5

Observation 491af39e-5970-49fd-a154-f335659969fa · inbound

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning cites this paper.

Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 68

Resolution
unresolved
no resolver link, observed 2026-07-12T05:48:27.255331Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T05:48:27.255331Z digest=sha256:536dc995e2c55eefbf3e343287e06f31b47b54d2c91223e33f255a42bef9a019

Observation 5e3d5295-6984-4f2b-80d0-a215eeeb8670 · inbound

TACG: Trajectory-Aware Commit Gating for Diffusion Language Model Decoding cites this paper.

TACG: Trajectory-Aware Commit Gating for Diffusion Language Model Decoding LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 46

Resolution
unresolved
no resolver link, observed 2026-07-12T03:56:26.770729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T03:56:26.770729Z digest=sha256:9e5bf652c9c5dc946ce11ae16afe217c23ae1b98396d83ee8bf6df61d2d23f4e

Observation 01287599-c0a9-451a-8b6d-7f3400a1fd62 · inbound

Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding cites this paper.

Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 36

Resolution
verified exact
local_arxiv, observed 2026-07-11T03:07:52.297701Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.

source=pdf_text observed=2026-07-11T03:04:12.500342Z digest=sha256:b5406fd3255d1432685fff95c3e801d7552c28f73a04b44b9305417aaeeb709b

Observation 29edb95b-2335-441a-86cd-ed5e4a95e833 · inbound

Discrete Diffusion Models: A Unified Framework from Tokenization to Generation cites this paper.

Discrete Diffusion Models: A Unified Framework from Tokenization to Generation LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 196

Resolution
unresolved
no resolver link, observed 2026-08-02T05:17:48.734640Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T05:17:48.734640Z digest=sha256:78e40d6ad739c2b408d0a0cf3dd62d9fa2ba5f0783bfd8dcddc632be45d0c5a4

Observation e2d08504-8703-4a64-b2f9-3723b98fc4bb · inbound

Polestar: Drift-Aware Cache Calibration and Token Commitment for Efficient Inference of Diffusion LLMs cites this paper.

Polestar: Drift-Aware Cache Calibration and Token Commitment for Efficient Inference of Diffusion LLMs LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-02T14:46:39.695123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T14:46:39.695123Z digest=sha256:417abe3679d51fe9d003b5b8a051e1a93f2642f89ea5c8cf37d3d31ffd1d9a4c

Observation 4151bba2-1a48-4929-8f1e-88859952e015 · inbound

Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation cites this paper.

Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-02T01:48:56.739924Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:48:56.739924Z digest=sha256:164b6f9812a081102d913498286aeb588275670618ffc1b10a25b7a7e18e91bf

Observation effdeb4d-4795-4f84-bf2d-209671065be8 · inbound

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding cites this paper.

ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-01T16:48:36.676187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T16:48:36.676187Z digest=sha256:9859719abc02e2d62c4291a3338ced71233425db864d1217bb882ebeb05980d7

Observation f5a08073-03ad-47be-96d7-8fc06ab571d3 · inbound

Faster but Different: Diagnosing and Controlling Content Drift in Accelerated Multimodal Diffusion Language Models cites this paper.

Faster but Different: Diagnosing and Controlling Content Drift in Accelerated Multimodal Diffusion Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T14:08:23.176270Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T14:08:23.176270Z digest=sha256:8dbb7a4fb9e81ead10f4c00ed270f1be2dc39e553d9e06620ff0ddb5edea126e

Observation 1479906e-e069-4285-84e4-0955a62c6fc0 · inbound

WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA cites this paper.

WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-06T00:40:52.308647Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T00:40:52.308647Z digest=sha256:3324bb717229b762a8a72a63e7562fb2d6a129414fc0315c4b2dd3af4b05eded

Observation 0764bdfc-f9c4-48fc-8f13-9bbc35667233 · inbound

Does More Retrieved Evidence Help Visual Retrieval-Augmented Generation with Diffusion Language Models? cites this paper.

Does More Retrieved Evidence Help Visual Retrieval-Augmented Generation with Diffusion Language Models? LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-10T16:51:18.515307Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-10T16:51:18.515307Z digest=sha256:aca252d10e31b2a1270b999f768ff9d37058c49d0c6d7418a282024b6ee7c407