Pith. sign in

Paper Citation Record · LEDGER

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models

As of 6 August 2026, this Paper Citation Record lists 40 of 40 outbound references and 2 inbound Pith citation observations for arXiv:2512.19433.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.19433 v2

Coverage vector

measured 40 of 40 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-16T20:38:06.225705Z

measured 42 of 42 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-06T06:34:29.942622+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T21:18:03.005508Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-19T21:22:48.470135Z

Reference resolution

40 of 40 outbound references displayed

  • verified exact20
  • verified fuzzy18
  • unresolved2
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 65f5083a-8cc3-4619-98b1-07463d790aff · outbound

This paper cites Johnson, Jonathan Ho, Daniel Tar- low, and Rianne van den Berg.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Johnson, Jonathan Ho, Daniel Tar- low, and Rianne van den Berg

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:38:25.218816Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:7a87022f8e157e7aa7a364e721a616d0df5b6bdafc6d4b12422ac0a044e98fdb

Observation f0d06e84-05ff-47fb-a5c1-77c3906aab20 · outbound

This paper cites an unresolved cited work.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Unresolved cited work

Reference 2

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:38:25.216527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:42fd7231bf84940468f40c4cd37ff88b3e33edfceea07472aab9ae6ffa06a8d6

Observation dc78bb8e-a098-4d54-8d4e-29d4a1367c72 · outbound

This paper cites Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Janus-Pro: Unified Multimodal Understanding and Generation with Data and Model Scaling

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:38:24.560495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:9e7d6579b89a439a845563dc5247e9f34f5235518606aeb54a6dd9444667dd07

Observation d54d9ef1-a8d1-4b9d-b1e1-483c5b322e30 · outbound

This paper cites Tts-var: A test- time scaling framework for visual auto-regressive genera- tion.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Tts-var: A test- time scaling framework for visual auto-regressive genera- tion

Reference 4

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:38:24.564028Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:53f89facddd0fc5c781710b23c857657c5f4bd011efb943b02047c4f2b7857da

Observation e7f7812b-587e-4c5c-aa3e-2f662e470395 · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Emerging Properties in Unified Multimodal Pretraining

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:38:24.557363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:adbf947cd45ab33064cfc2475e0ca8126378001e407f1e490ebbd81cc50cafd1

Observation bef2c71e-1534-4a91-9094-ab99cce6dc01 · outbound

This paper cites Lumina-t2x: Scalable flow-based large diffusion transformer for flexible resolution generation.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Lumina-t2x: Scalable flow-based large diffusion transformer for flexible resolution generation

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:38:25.214301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:21a68e4d869a9b62146beeeda0d0ccecb6b828e01162a610d10738213fc32af5

Observation 5f206ea4-c32b-4670-82ee-095090eed136 · outbound

This paper cites Geneval: An object-focused framework for evaluating text- to-image alignment.Advances in Neural Information Pro- cessing Systems (NeurIPS), 36.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Geneval: An object-focused framework for evaluating text- to-image alignment.Advances in Neural Information Pro- cessing Systems (NeurIPS), 36

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:38:25.188186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:f0eefbc8cc0e6dcc18f67362c615787f255ea825d785a1a969c46f0866891d22

Observation 0fa3ca06-6f35-4228-8010-5901375f39f0 · outbound

This paper cites Scaling diffusion language models via adap- tation from autoregressive models.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Scaling diffusion language models via adap- tation from autoregressive models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:38:25.192653Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:fb85c2c06cdb93b74cb04c62942919e0a75605b3d13cfe60866223fe5743d777

Observation 736b3d41-491a-4ba5-9cbc-87304f63e834 · outbound

This paper cites Clipscore: A reference-free evaluation met- ric for image captioning.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Clipscore: A reference-free evaluation met- ric for image captioning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:38:25.179999Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:bb87bcf6bd7383d9e6d2f01e8f58df234af18260a83dcde5e1d16e8c516a8725

Observation 0ded1d3e-bfd9-4f7c-a529-6e05cc0f00a5 · outbound

This paper cites Denoising diffu- sion probabilistic models.Advances in Neural Information Processing Systems (NeurIPS).

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Denoising diffu- sion probabilistic models.Advances in Neural Information Processing Systems (NeurIPS)

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:38:25.182205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:a4ba4c712cb1ae2a8b77058c3697d0aa6d1f1e5e1e19f96be73b567fdbc36fbe

Observation ab0c33f7-4694-4547-87d2-b120fe682341 · outbound

This paper cites GPT-4o System Card.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models GPT-4o System Card

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:38:24.500559Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:a84e2512ab948efe3c24172170bab051388aa74e13c82cf4607989426b192ffe

Observation 8677a576-4fde-4268-aacf-b4747b44ded0 · outbound

This paper cites OpenAI o1 System Card.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models OpenAI o1 System Card

Reference 12

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:38:24.524386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:ded15e30021890e1ab8e44ac3fbc3e8aab00a22b8d995be0a57e7e547294c06c

Observation e3bf41a0-b568-45ca-b818-ff27c18376d0 · outbound

This paper cites Flux.https://github.com/ black-forest-labs/flux.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Flux.https://github.com/ black-forest-labs/flux

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:38:25.200679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:c62e7dc4567226665e94cd7fdc37a11b65002b8ba72d365419913c1aa74747fe

Observation e5c493a0-3d7d-4841-9d6b-733ff71b116d · outbound

This paper cites Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Reflect-DiT: Inference-Time Scaling for Text-to-Image Diffusion Transformers via In-Context Reflection

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:38:24.507584Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:7763c54b199d181a7bd721b0fa913d2abe3c6894a1912393b8a65de957956d74

Observation 5d334b33-de6f-4a3d-a3ab-566b9c84a901 · outbound

This paper cites Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Lumina-mGPT: Illuminate Flexible Photorealistic Text-to-Image Generation with Multimodal Generative Pretraining

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:38:24.554015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:9120d967b558d50942bdbe09799bd70e26887ce61295d71f0508378907095586

Observation c59834e0-a743-419d-a13a-721c106db13f · outbound

This paper cites Discrete diffusion modeling by estimating the ratios of the data dis- tribution.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Discrete diffusion modeling by estimating the ratios of the data dis- tribution

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:38:25.177747Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:2b9b8d0b630014193ce687f6a8cc789406c27647415ab6715dc2a7c0ac1d964d

Observation b45a99d3-1aee-4131-98e4-9024edb7b6bb · outbound

This paper cites Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Inference-Time Scaling for Diffusion Models beyond Scaling Denoising Steps

Reference 17

Resolution
verified exact
arxiv_id, observed 2026-05-20T11:45:17.785272Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:803fe3d6e0e6dd16a673e3d243c08e5e6dd0f5bc9cc178d44910c5a8a60acf6d

Observation 0bbee2b1-de59-4091-b187-3eb2c2dd58da · outbound

This paper cites Large Language Diffusion Models.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Large Language Diffusion Models

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:38:24.510754Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:56a737574c7fb10ed5bd5d61175828ff043782360e1d67154360a50a8c57c8c0

Observation 18f50fff-142c-4059-b3e5-f8679a4d1435 · outbound

This paper cites aMUSEd: An Open MUSE Reproduction.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models aMUSEd: An Open MUSE Reproduction

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:38:24.550578Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:9292d9d33951d715955d97625714f0ad8c6e6afa77af15452422f1a571d06058

Observation cf8dd3fc-8f7d-443f-badc-5e616039c7bd · outbound

This paper cites Sdxl: Improving latent diffusion models for high-resolution image synthesis.Proceedings of the In- ternational Conference on Learning Representations (ICLR).

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Sdxl: Improving latent diffusion models for high-resolution image synthesis.Proceedings of the In- ternational Conference on Learning Representations (ICLR)

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:38:25.190325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:aced4661055cc3b264a792055f6a870c0c27a81228ce0144c4ac19bb956a690f

Observation 299defab-a04a-4204-8521-30bcfec9dc45 · outbound

This paper cites Lumina- image 2.0: A unified and efficient image generative frame- work.Proceedings of the IEEE International Conference on Computer Vision (ICCV).

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Lumina- image 2.0: A unified and efficient image generative frame- work.Proceedings of the IEEE International Conference on Computer Vision (ICCV)

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:38:25.184303Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:c2822aa68c63bc34361629b6c746fc0594d05db14be721d163a4685f0dae8afc

Observation 83f506e9-11ab-46b3-9a90-0ef89805506f · outbound

This paper cites Learn- ing transferable visual models from natural language super- vision.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Learn- ing transferable visual models from natural language super- vision

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:38:25.186231Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:1600e5b0a9eb7dd3ca0dac70e5dd92b4de0ff4f939f527c41efa8c96381c4705

Observation ed377402-831b-407a-b462-7f89a393c3f5 · outbound

This paper cites High-resolution image syn- thesis with latent diffusion models.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models High-resolution image syn- thesis with latent diffusion models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:38:25.207871Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:a4065bb2bbce6e9a32b31a186deaaaa1055055420df697cace7f53c2b2f5f61d

Observation 64a654cf-6ac5-45ed-bd58-20b7682bfd5f · outbound

This paper cites Muddit: Liber- ating generation beyond text-to-image with a unified discrete diffusion model.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Muddit: Liber- ating generation beyond text-to-image with a unified discrete diffusion model

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:38:25.210148Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:4e334944d648d7cef4878fddb286177041827eb39ff7b8c31d175a17799e7fd3

Observation 00180444-fc9d-4874-bdad-d1a9920a0164 · outbound

This paper cites A General Framework for Inference-time Scaling and Steering of Diffusion Models.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models A General Framework for Inference-time Scaling and Steering of Diffusion Models

Reference 25

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:38:24.514296Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:1f2c48839048efe07db471f3ca10722c8dda14c49d3d162ba275182a6f20e672

Observation ef16b56f-5f92-41a7-b955-41f8835061f5 · outbound

This paper cites Emu3: Next-Token Prediction is All You Need.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Emu3: Next-Token Prediction is All You Need

Reference 26

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:38:24.543869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:11dc090be083ceff4be8bf6cb7172cee4cca8af63ba596eda3c6cc7411c12c19

Observation d818f328-9552-44a1-8ae5-a5ff1c164246 · outbound

This paper cites Qwen-Image Technical Report.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Qwen-Image Technical Report

Reference 27

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:38:24.540825Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:e3ce4f2efb457e62f08918681d1a226d7257ddbec650e4543bd491d2790c4acb

Observation 37dbb905-08c6-4e00-8201-2cb9c723509f · outbound

This paper cites Sana: Efficient high- resolution image synthesis with linear diffusion transform- ers.Proceedings of the International Conference on Learn- ing Representations (ICLR).

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Sana: Efficient high- resolution image synthesis with linear diffusion transform- ers.Proceedings of the International Conference on Learn- ing Representations (ICLR)

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:38:25.202875Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:28e705ee55d482012fe8d6800d87ebd46fb8418ea8be472e9e52bc6fe5217215

Observation 1b870263-c76d-45d5-ac1f-9551aac673b3 · outbound

This paper cites Sana 1.5: Efficient scaling of training-time and inference-time compute in linear diffusion transformer.Proceedings of the International Conference on Machine Learning (ICML).

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Sana 1.5: Efficient scaling of training-time and inference-time compute in linear diffusion transformer.Proceedings of the International Conference on Machine Learning (ICML)

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:38:25.205472Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:be475e6e059c4d5940db30c1225e6b184efa7e66c5028490b2ad1d35319061d8

Observation d3351bb1-0e6e-47a4-a585-4f62a17adb61 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 30

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:38:24.547194Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:dd069dd9dc5be62991a9ce29aae251bcc927316942798bfdfb9e549aa2713bdb

Observation 512da9c3-4cf3-4f78-acf1-f91c881879b9 · outbound

This paper cites Lumina-dimoo: An omni diffusion large language model for multi-modal generation and understanding.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Lumina-dimoo: An omni diffusion large language model for multi-modal generation and understanding

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:38:24.537938Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:16f6604f8664469f7e7feb31c8f0fa49b6688ebeb84af7c0f334023f307a8ddb

Observation feb673e6-274a-4997-8d09-7064a73bb1e4 · outbound

This paper cites Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Lumina-mGPT 2.0: Stand-Alone AutoRegressive Image Modeling

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:38:24.534551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:eeb7e6cfbaf7f8f11e0477e3bbf1a49d9781fff2c6aa4f1d4478f7e74ee7d05e

Observation 1310161c-f795-4554-b47c-2c561285cc9a · outbound

This paper cites Resurrect Mask AutoRegressive Modeling for Efficient and Scalable Image Generation.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Resurrect Mask AutoRegressive Modeling for Efficient and Scalable Image Generation

Reference 33

Resolution
verified exact
arxiv_id, observed 2026-05-16T20:38:24.531104Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:5a3a29d61f2f952d4b004e94288cd9e2e73c1fc7e73ee7c12658b5b132827cfd

Observation ce9cb3b7-1c9e-45dd-a978-b8b10e7a0bd6 · outbound

This paper cites Imagere- ward: Learning and evaluating human preferences for text- to-image generation.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Imagere- ward: Learning and evaluating human preferences for text- to-image generation

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:38:25.198856Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:a6a387eab458c1b38dd3b85a106eb477eb8578b522c1c258a1e23e373ae7a050

Observation df090975-d48e-4dbf-acf9-6cc62755331b · outbound

This paper cites MMaDA: Multimodal Large Diffusion Language Models.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models MMaDA: Multimodal Large Diffusion Language Models

Reference 35

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:38:24.504106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:1095eed1720becfab85bdacc3687128628224615cba7203ec18b190a2f55c16c

Observation 41275aa0-15dd-4336-9322-7128ac8da6ef · outbound

This paper cites Towards understanding the working mechanism of text-to-image dif- fusion model.Advances in Neural Information Processing Systems (NeurIPS).

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Towards understanding the working mechanism of text-to-image dif- fusion model.Advances in Neural Information Processing Systems (NeurIPS)

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:38:25.212202Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:af16b8d5f776aebb57607f6bd45f1ae4606e6ac2e0cbd8a9e5545e080a110fbf

Observation 7312622e-1578-4eea-aaf4-5a9bc24a6733 · outbound

This paper cites LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 37

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:46:06.580274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:1f899e11ae34268d5fbb1c28495dbc37725df4cec302dc97b0298c05b63a6fb3

Observation c5e3fe73-6cd6-4814-a38f-2d829c046c5c · outbound

This paper cites LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models LLaDA 1.5: Variance-Reduced Preference Optimization for Large Language Diffusion Models

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-16T20:38:24.527750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:e02b5c00bae9b29b386bdf805d0dd4161eb0ede4e154ec258a03b4e03e6575c1

Observation 6c8fcc9e-7f6b-4ad3-b5e4-e11522056893 · outbound

This paper cites Lumina-next: Making lumina-t2x stronger and faster with next-dit.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Lumina-next: Making lumina-t2x stronger and faster with next-dit

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-05-16T20:38:25.194741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:28b7f7b1ccb2fd08c205b7bf2c4363a2dccede9a9faba7d7d27992345c786713

Observation 9480c05b-fc94-4d75-ac0b-34f3e1e8a593 · outbound

This paper cites an unresolved cited work.

dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models Unresolved cited work

Reference 40

Resolution
unresolved
raw_fallback, observed 2026-05-16T20:38:25.196671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-16T20:38:06.225705Z digest=sha256:144915f4a59f0550f472b88a294c89af8f58f4653292bec5296484b874efff5c

Pith citing papers

Observation cff2cb2e-f720-4f6e-9aaa-13989a23adab · inbound

Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model cites this paper.

Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-05-19T13:02:18.391456Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T12:59:31.454155Z digest=sha256:8af64992255ea24c6349120d963d6b93ea79b66d1cbcad7ff74ef64a82766d73

Observation 21d20a71-dd7b-4be5-8d7f-261f920d7e64 · inbound

Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models cites this paper.

Sketch Then Paint: Hierarchical Reinforcement Learning for Diffusion Multi-Modal Large Language Models dMLLM-TTS: Self-Verified and Efficient Test-Time Scaling for Diffusion Multi-Modal Large Language Models

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-05-19T21:22:48.472803Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-06T06:34:29.942622+00:00.

source=pdf_text observed=2026-05-19T21:18:03.005508Z digest=sha256:e6ef6a6033ad28c1ff4ae38f73424a630991cc72b36132b040d7c84460f7e154