Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T03:46:06.074416Z
Paper Citation Record · LEDGER
As of 12 August 2026, this Paper Citation Record lists 100 of 125 outbound references and 59 inbound Pith citation observations for arXiv:2505.16933.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T03:46:06.074416Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-12T06:34:41.77262+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T16:51:18.515307Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-11T03:07:52.266883Z
100 of 125 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation eb3827e3-a287-405e-a693-ca44886b26e5 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Visual instruction tuning
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 1da84e56-3bd4-4ebe-ab82-4a7ef24e5818 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Improved baselines with visual instruction tuning
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3e14fcb8-15a1-4822-85bd-75e105605ac1 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning LLaVA-OneVision: Easy Visual Task Transfer
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f169e065-3353-4bc7-8af7-153c59d0da1b · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d0c19dc5-0c8a-4ca8-9786-8f45aa59691a · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4d8324a0-973b-47e8-b9cd-e6b79cde81f7 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Chameleon: Mixed-Modal Early-Fusion Foundation Models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4ce34fd3-d640-4d04-896e-3421ccbaa926 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Kimi-Audio Technical Report
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8b25ced2-3437-4776-af5b-64709d4dc2ab · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Qwen-Audio: Advancing Universal Audio Understanding via Unified Large-Scale Audio-Language Models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 06d1d233-677c-4da6-a2ad-61236a8de789 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning GAMA: A Large Audio-Language Model with Advanced Audio Understanding and Complex Reasoning Abilities
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 40cb993f-0961-4de2-82ad-085c8318420d · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning InternVideo2.5: Empowering Video MLLMs with Long and Rich Context Modeling
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 166dcc9d-cf00-4591-be91-8ac64bfe4c22 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Sharegpt4video: Improving video understanding and generation with better captions
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 331ed5d7-2084-415c-bfeb-bc831b564724 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning LLaVA-Video: Video Instruction Tuning With Synthetic Data
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b7a15082-b335-4343-b195-3bde33bcc2d9 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Improving language understand- ing by generative pre-training
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bcba5c8f-c61c-4fd9-a7d7-35d5f3be3d9c · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Language models are unsupervised multitask learners
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 48d415b4-551d-40c4-aef8-c91e1e3283ad · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Language models are few-shot learners
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d9e3779f-853a-401e-92bc-4b97125945e0 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning LLaMA: Open and Efficient Foundation Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9f4fe500-7a64-4a75-a2a4-354d104ff9bd · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Llama 2: Open Foundation and Fine-Tuned Chat Models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 863a75d8-8b69-4ae0-8d25-7ca290de2a65 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning The Llama 3 Herd of Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9ea18293-1c8d-4055-b5df-92e8a040a93b · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Qwen2.5 Technical Report
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9799e762-59f7-4ad2-9188-2c2e8e84faa2 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Textbooks Are All You Need II: phi-1.5 technical report
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e20982b7-90c8-46b3-8f8a-598f0a197b53 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning DeepSeek LLM: Scaling Open-Source Language Models with Longtermism
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9d44a0bf-8555-41d1-9c94-7eee55cd6dfe · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Deep unsupervised learning using nonequilibrium thermodynamics
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e804b2df-89ef-4240-88a4-325bfb7d01da · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Denoising diffusion probabilistic models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 34e3f717-f466-4afe-9689-add636dc0769 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Score-Based Generative Modeling through Stochastic Differential Equations
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8edaaf77-deae-4cfd-8618-38e6b417e85a · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Argmax flows and multinomial diffusion: Learning categorical distributions
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7c97355f-9f1e-43b8-ac90-3c6554c5ba9f · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Structured denoising diffusion models in discrete state-spaces
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f52b048d-4123-4f13-83c6-bedd5d61d168 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning One transformer fits all distributions in multi-modal diffusion at scale
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0223dedc-50b9-4d17-ab40-db591575af9e · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Show-o: One Single Transformer to Unify Multimodal Understanding and Generation
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 296627a3-06a6-4c75-89c3-b84188b0c60d · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f69b31da-aeac-4421-b8fc-b816f8812f13 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning JanusFlow: Harmonizing Autoregression and Rectified Flow for Unified Multimodal Understanding and Generation
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 839ae744-19dc-4705-a683-de022f9a29e4 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning MetaMorph: Multimodal Understanding and Generation via Instruction Tuning
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d0d52794-3986-4ec9-92e2-92310a2bf8d1 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Orthus: Autoregressive Interleaved Image-Text Generation with Modality-Specific Heads
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3f4cc231-4116-41ba-a654-3b56db6a9fd6 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Unified Multimodal Discrete Diffusion
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b40dcffc-9bae-4483-bd3b-485d13ea5e07 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Dual Diffusion for Unified Image Generation and Understanding
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation da41ae7d-9522-4068-a4ad-eb02fdbf6299 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning A continuous time framework for discrete denoising models
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d5937ee9-e1f8-45ff-aa30-2c14e0817adb · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning DiffusionBERT: Improving Generative Masked Language Models with Diffusion Models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5568ff17-9370-4799-ab26-0f87be51fb50 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Score-based continuous-time discrete diffusion models
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d00570c7-a48a-45eb-9999-4f9bb61511b4 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Discrete diffusion modeling by estimating the ratios of the data distribution
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c524b596-ef42-44fe-9b85-ed2d972577e3 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Simplified and Generalized Masked Diffusion for Discrete Data
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e7475394-6ea5-4e0c-9ed7-246eb2dc315c · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Simple and Effective Masked Diffusion Language Models
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6acf96e2-3ce4-45cc-9a1e-10feeba17223 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Your Absorbing Discrete Diffusion Secretly Models the Conditional Distributions of Clean Data
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9bd7c7d8-9e2d-4f4a-8089-1341606f43fb · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Large Language Diffusion Models
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bd0b8896-2c9f-4690-baac-3ec50bfdcd6c · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Effective and efficient masked image generation models
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0c0d6acd-f0e8-4f5c-b306-bf8ba73ece36 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cc9b7f28-f032-498c-b772-ba0ac2edc3dc · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 445b6210-ae1a-4947-bae5-b525d1df4109 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Scaling up Masked Diffusion Models on Text
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 28decc83-d2eb-411a-b305-0dbd6b0f38be · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Generative flows on discrete state-spaces: Enabling multimodal flows with applications to protein co-design
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 08656729-0c85-4f49-a871-8b13c4e11075 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning [MASK] is All You Need
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 41c3e79d-43c6-47d2-a54b-8e3497ef3a44 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Learning transferable visual models from natural language supervi- sion
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 235af7ba-d005-4b4e-a10b-404e509d1973 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Sigmoid loss for language image pre- training
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f3554d0f-20eb-43be-b3ad-ef62fd9bae44 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Wan: Open and Advanced Large-Scale Video Generative Models
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b7916aac-0447-4d56-aceb-d18a2a28da93 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning CogVideoX: Text-to-Video Diffusion Models with An Expert Transformer
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 43fc7807-9237-4cae-b9f5-66485888a76e · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning HunyuanVideo: A Systematic Framework For Large Video Generative Models
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 97d3e00c-f044-47ea-95a8-d494f548746f · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Llava-next: Improved reasoning, ocr, and world knowledge
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 74420e64-e9bf-490e-9bad-8518905b5848 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 56107efc-5ba8-45f9-8914-68e61b75a389 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 766fedce-a1e5-4383-85b2-bef96fdbbfe2 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Qwen3: Think deeper, act faster
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b09d0fa8-ad5f-4e5d-b554-7ded94b6979a · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Proximal Policy Optimization Algorithms
Reference 58
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e69c9a60-9b6b-46eb-a1da-eaf0754f3d8d · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Direct pref- erence optimization: Your language model is secretly a reward model
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ae8df455-53cf-436d-b6a1-2d4c87990ae6 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Simpo: Simple preference optimization with a reference-free reward
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 25e23fe3-dd3c-4c96-8eb7-124c01a569c0 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d357fa1e-e6c9-4c78-9844-a96a9eac384f · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 59aa0ef3-eed0-4f7e-870d-2f320e92a188 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark
Reference 63
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 954aae21-7be7-4b15-811b-87e85925fe7b · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 64
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 96b9674b-7c79-4d0c-8582-4f29ae20bde2 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 65
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 94533f71-da6b-4817-9cd0-cfd613a27ac5 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Mmbench: Is your multi-modal model an all-around player?
Reference 66
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6d754cfa-d9d4-4afd-adff-19a806badf15 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems?
Reference 67
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0dfe3396-9f15-45bb-a736-7942eb673620 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Mathvista: Evaluating math reasoning in visual contexts with gpt-4v, bard, and other large multimodal models
Reference 68
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9dd7dea5-c8ce-406a-a6f2-67c4e0c9f396 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning A diagram is worth a dozen images
Reference 69
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f6c0efea-4f43-47c8-a549-b2bac16b1c8a · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a32c1cb8-8f25-42d2-817d-2bf52fdc48e3 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Docvqa: A dataset for vqa on document images
Reference 71
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 097ac91d-3457-4070-9e4e-facd08cb09ef · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Infographicvqa
Reference 72
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d4403c52-757d-4631-9544-b237b722ab18 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Grok-1.5 vision preview
Reference 73
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 546a3dee-5b6d-4675-a23e-1e94049a767d · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
Reference 74
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 34eaec19-4036-40b3-813d-b147c32d8412 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning MLVU: Benchmarking Multi-task Long Video Understanding
Reference 75
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 4362463a-c272-4aa9-9fa4-c71980ac5c6c · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Video-MME: The First-Ever Comprehensive Evaluation Benchmark of Multi-modal LLMs in Video Analysis
Reference 76
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 079da712-36c7-4575-8ad6-f9906f60615e · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Sharegpt4v: Improving large multi-modal models with better captions
Reference 77
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 98b271ca-217a-457f-99b8-2aae51ed6882 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Cambrian-1: A fully open, vision-centric exploration of multimodal llms
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f926c3c7-f79c-4371-a099-c6908596a461 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning DeepSeek-VL: Towards Real-World Vision-Language Understanding
Reference 79
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 49a6cef4-ddf7-432d-80ca-9b727dc6395c · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning DeepSeek-VL2: Mixture-of-Experts Vision-Language Models for Advanced Multimodal Understanding
Reference 80
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 010a517e-f0b1-4adc-9d7c-92ee868653fc · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Janus: Decoupling Visual Encoding for Unified Multimodal Understanding and Generation
Reference 81
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 415cfe4b-de35-47ff-9419-73ac225adc98 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d483a297-5ab9-499c-b3ed-e7815f5a501e · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Emu3: Next-Token Prediction is All You Need
Reference 83
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 90c34c26-3cb2-43a8-9c6f-4cb44e5ec9cd · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning What matters when building vision- language models?
Reference 84
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6335c45f-b843-448d-85f7-16fa449c88e4 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Diffusion-lm improves controllable text generation
Reference 85
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 06e83cfe-da1b-4941-b442-33935c749004 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning DiffuSeq: Sequence to Sequence Text Generation with Diffusion Models
Reference 86
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b375bf3c-173b-4363-ab3b-cb69fb510164 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning SSD-LM: Semi-autoregressive Simplex-based Diffusion Language Model for Text Generation and Modular Control
Reference 87
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 5e5f35e3-e498-4783-88d0-c5f65f595e2b · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Self-conditioned Embedding Diffusion for Text Generation
Reference 88
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 826ada18-bbca-486e-abb0-90a06845a8e1 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Analog Bits: Generating Discrete Data using Diffusion Models with Self-Conditioning
Reference 89
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation aee58e8e-e87d-46ec-b43d-2163b6a6cd53 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Continuous diffusion for categorical data
Reference 90
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 335ae4b8-e1ab-445d-8b78-1d88a7ac3d12 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Categorical sdes with simplex diffusion
Reference 91
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation fc0cb21d-4572-43b1-9348-b787ecbec784 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Ar-diffusion: Auto-regressive diffusion model for text generation
Reference 92
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 79705010-b49f-4211-9800-40a077a77b6b · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Tess: Text-to-text self-conditioned simplex diffusion
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9f78f635-7cfe-4958-9aa9-9718fecfd85a · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning DINOISER: Diffused Conditional Sequence Learning by Manipulating Noises
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 89cc55f2-661b-454f-b7c3-701027ffb29c · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Planner: Generating diversified paragraph via latent language diffusion model
Reference 95
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation cd448a65-09ea-4803-9e1f-529f3b07ec74 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Reflected diffusion models
Reference 96
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b78038ef-7bf1-429d-b8f5-e0cc015bc126 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Bayesian Flow Networks
Reference 97
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 2f7b1ba1-11c5-47ea-ba0a-1e63ff2b25f8 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Text generation with diffusion language models: A pre-training approach with continuous paragraph denoise
Reference 98
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 82d4db9d-1bb7-442d-a71c-07dc8d9c0f1e · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Unifying Bayesian Flow Networks and Diffusion Models through Stochastic Differential Equations
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c40fe4d7-3a34-4b71-927d-6fb9afbb34c5 · outbound
LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning Target Concrete Score Matching: A Holistic Framework for Discrete Diffusion
Reference 100
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c7785dfa-a5d3-4648-ae12-9fccc85b42cb · inbound
Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation a82829d2-a820-4e1e-a799-5096c180d32f · inbound
The Philosophy and Physics of Duality LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 07b2dcdb-1cff-428a-9a8c-6e1035ba0070 · inbound
A Survey on Diffusion Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 072f861c-8656-4bc4-935d-2ad3398040cc · inbound
LLaDA-VLA: Vision Language Diffusion Action Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dca79717-210d-479b-90c9-37e24c8d55fc · inbound
Inpainting-Guided Policy Optimization for Diffusion Large Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 24b21a0c-3418-40c6-8124-7baa99028163 · inbound
Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 71
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40f9660a-6be0-4cce-93c0-e954c9b8ac14 · inbound
CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Credit LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 02d7388a-13b4-4c4f-b1fd-67961e2f98f9 · inbound
CreditDecoding: Accelerating Parallel Decoding in Diffusion Large Language Models with Trace Credit LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 87f4633e-4417-412e-8e99-a5245e57c9db · inbound
A Comprehensive Study on Visual Token Redundancy for Discrete Diffusion-based Multimodal Large Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation de50c696-4650-4c37-a92b-b230fba37457 · inbound
Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 69
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cb444cf3-6c48-4bd9-ac63-0a6dd390f5df · inbound
Efficient-DLM: From Autoregressive to Diffusion Language Models, and Beyond in Speed LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 7312622e-1578-4eea-aaf4-5a9bc24a6733 · inbound
dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 91116f92-a9f5-4328-b953-eb8627c20fa7 · inbound
Streaming-dLLM: Accelerating Diffusion LLMs via Suffix Pruning and Dynamic Decoding LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36b693b4-2464-4076-a5c0-d10994d5788e · inbound
DODO: Discrete OCR Diffusion Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1af74f9a-ebed-46be-bb48-014677dea550 · inbound
Omni-Diffusion: Unified Multimodal Understanding and Generation with Masked Discrete Diffusion LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 92a023c9-53c7-496f-a97c-b996afba0642 · inbound
Thinking Diffusion: Penalize and Guide Visual-Grounded Reasoning in Diffusion Multimodal Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d5b250fd-1231-4206-9f3b-fbbad600a124 · inbound
DMax: Aggressive Parallel Decoding for dLLMs LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 48172090-4760-4a73-96a6-c70627854ba7 · inbound
DMax: Aggressive Parallel Decoding for dLLMs LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 94
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 6ebb8afe-175f-4c15-93df-684a05f7cbbc · inbound
ECHO: Efficient Chest X-ray Report Generation with One-step Block Diffusion LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 401f5876-6c26-4a02-85cf-c91df8f2ce99 · inbound
ECHO: Efficient Chest X-ray Report Generation with One-step Block Diffusion LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 52b44d06-1538-478c-8960-9592f83ddfaf · inbound
LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 56
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 10beb142-c2e7-4b3f-b268-6d919da7e65c · inbound
LaDA-Band: Language Diffusion Models for Vocal-to-Accompaniment Generation LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 53
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1686a21a-cb41-41e2-9221-edc480f3f1df · inbound
Dataset-Level Metrics Attenuate Non-Determinism: A Fine-Grained Non-Determinism Evaluation in Diffusion Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c15f2d8e-0392-4f8f-84ee-744cf0776978 · inbound
DF3DV-1K: A Large-Scale Dataset and Benchmark for Distractor-Free Novel View Synthesis LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fbbafb58-abfd-4d16-b8ca-a2b898647d8e · inbound
BARD: Bridging AutoRegressive and Diffusion Vision-Language Models Via Highly Efficient Progressive Block Merging and Stage-Wise Distillation LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation eb4bde6f-c13e-46f7-a290-c7abca04ea42 · inbound
Stability-Weighted Decoding for Diffusion Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ca8d6825-d1c7-47a7-aaf5-00c0fd0daf9a · inbound
One Step Forward and K Steps Back: Better Reasoning with Denoising Recursion Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 173
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0edd5724-900e-418f-9547-dc81ff4795ad · inbound
Continuous Latent Diffusion Language Model LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 105
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation b34a8d5b-8b14-4ff0-8664-9ddca8b66942 · inbound
GPO-V: Jailbreak Diffusion Vision Language Model by Global Probability Optimization LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c57969a2-7928-4a43-bffd-f52edbfd67af · inbound
GPO-V: Jailbreak Diffusion Vision Language Model by Global Probability Optimization LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 3a02f2a6-a770-466f-81a8-1cb2f278de4c · inbound
Relative Score Policy Optimization for Diffusion Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 99
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9a8fe905-5117-403b-91d0-fdf1754ade04 · inbound
ELF: Embedded Language Flows LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f89f8dc6-cc6c-4f68-9a6f-f28f0e3c50d5 · inbound
ELF: Embedded Language Flows LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 78
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 0541e5a3-d182-4862-882c-223458aa6288 · inbound
MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation c046c1a7-c9b0-4ecd-b767-f7c846f2431f · inbound
MindVLA-U1: VLA Beats VA with Unified Streaming Architecture for Autonomous Driving LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 8a319274-e01d-4b42-b589-99229e3338c1 · inbound
Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d91d7967-7d2b-4495-9d9c-4fc7050ec5a9 · inbound
Mitigating Mask Prior Drift and Positional Attention Collapse in Large Diffusion Vision-Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 97d16801-5838-4526-827a-d903d58cf0e7 · inbound
Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation bf27191e-27c8-476a-a3fc-64e88f6847c7 · inbound
Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 941648a2-a284-49f5-8c8a-8a06e1b960ee · inbound
Fast-dDrive: Efficient Block-Diffusion VLM for Autonomous Driving LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation e3c9f15a-40bb-4901-940c-6be7a95ad9c1 · inbound
Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation f822b833-a84e-436b-b1f1-f51754cfbdf7 · inbound
AnyMo: Scaling Any-Modality Conditional Motion Generation with Masked Modeling LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 9ae228eb-72a0-4e85-a531-164d1de881bb · inbound
dMoE: dLLMs with Learnable Block Experts LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation ff7ad414-ffcb-43a5-b5cd-3bb6c841210e · inbound
Dynamic Infilling Anchors for Format-Constrained Generation in Diffusion Large Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 82
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 930ae589-b850-42be-8cc0-a8501c1e2886 · inbound
Prefilling-dLLM: Predictive Prefilling for Long-Context Inference in Diffusion Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation d13cc417-d2c5-4d66-8cfd-cc2f299eca4d · inbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 32551ffa-a4de-414c-8aa5-6147c706dc9e · inbound
Improved Large Language Diffusion Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 02c0a308-8bee-444f-8aa6-fc63c7094570 · inbound
Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 09b26b26-9119-418d-bf41-2f6e9082c521 · inbound
Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 52
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 491af39e-5970-49fd-a154-f335659969fa · inbound
Parallelized Autoregressive Decoding for Omni-Modal Dense Video Captioning LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 68
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5e3d5295-6984-4f2b-80d0-a215eeeb8670 · inbound
TACG: Trajectory-Aware Commit Gating for Diffusion Language Model Decoding LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01287599-c0a9-451a-8b6d-7f3400a1fd62 · inbound
Nemotron-Labs-Diffusion: A Tri-Mode Language Model Unifying Autoregressive, Diffusion, and Self-Speculation Decoding LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-12T06:34:41.77262+00:00.
Observation 29edb95b-2335-441a-86cd-ed5e4a95e833 · inbound
Discrete Diffusion Models: A Unified Framework from Tokenization to Generation LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 196
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e2d08504-8703-4a64-b2f9-3723b98fc4bb · inbound
Polestar: Drift-Aware Cache Calibration and Token Commitment for Efficient Inference of Diffusion LLMs LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4151bba2-1a48-4929-8f1e-88859952e015 · inbound
Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 73
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation effdeb4d-4795-4f84-bf2d-209671065be8 · inbound
ST-Veto: Spatio-Temporal Token Veto for Diffusion MLLMs via Taylor Prediction and Visual Grounding LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5a08073-03ad-47be-96d7-8fc06ab571d3 · inbound
Faster but Different: Diagnosing and Controlling Content Drift in Accelerated Multimodal Diffusion Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1479906e-e069-4285-84e4-0955a62c6fc0 · inbound
WAM-Diff2: Hierarchical AR-to-Diffusion Distillation for Highly Efficient Autonomous Driving VLA LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0764bdfc-f9c4-48fc-8f13-9bbc35667233 · inbound
Does More Retrieved Evidence Help Visual Retrieval-Augmented Generation with Diffusion Language Models? LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.