Pith. sign in

Paper Citation Record · LEDGER

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models

As of 14 August 2026, this Paper Citation Record lists 82 of 82 outbound references and 5 inbound Pith citation observations for arXiv:2512.14008.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2512.14008 v2

Coverage vector

measured 82 of 82 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T16:21:40.277901Z

measured 87 of 87 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-14T06:32:32.682623+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T09:39:33.093629Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T17:38:43.676950Z

Reference resolution

82 of 82 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved81
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 2e5f1d29-7e7a-4fb7-b411-955f2a7ea27c · outbound

This paper cites Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Block Diffusion: Interpolating Between Autoregressive and Diffusion Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:31.275013Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:31.275013Z digest=sha256:cf0edaadcfe4dc84060cd7e42dd6bc348342d36d0ec23c96dbccae14e1c64fc8

Observation b8146d99-7772-4259-a545-b88a37ed3a2f · outbound

This paper cites Structured denoising dif- fusion models in discrete state-spaces.Advances in neural information processing systems, 34:17981–17993, 2021.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Structured denoising dif- fusion models in discrete state-spaces.Advances in neural information processing systems, 34:17981–17993, 2021

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:31.317490Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:31.317490Z digest=sha256:501da73a106ced25749df1696866988696dbebb13647c92a8766ed48542bb826

Observation 015f33cc-23d0-4e5c-9e3d-196760368038 · outbound

This paper cites Qwen2.5-VL Technical Report.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:31.374567Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:31.374567Z digest=sha256:56196699da85c276456ad4ace43d744eb6839e3d50c17c5a331bdb56751a747a

Observation d6221954-d337-4a12-b8fe-e52ebc046a27 · outbound

This paper cites Halton Scheduler For Masked Generative Image Transformer.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Halton Scheduler For Masked Generative Image Transformer

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:31.435181Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:31.435181Z digest=sha256:90630cd7f7aaa7e55c5657213606aec41b500901c172cf8ec7197f2741246015

Observation 83c143c3-c219-4a5a-9e2b-819970be1393 · outbound

This paper cites In- structpix2pix: Learning to follow image editing instructions.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models In- structpix2pix: Learning to follow image editing instructions

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:31.520363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:31.520363Z digest=sha256:b137ad186f196834cf5b88a7ae6a85002c22a91feb84bbb118536c6bf1a8b765

Observation dd5ac18c-1554-45e5-b067-298f16be34a3 · outbound

This paper cites Coyo-700m: Image-text pair dataset.https : / / github.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Coyo-700m: Image-text pair dataset.https : / / github

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:31.603156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:31.603156Z digest=sha256:a62904da2e20af34356b179655314cf06319fd936e32fd930ab1de87002d13c3

Observation 95c26570-d1c3-4547-b6b1-59553bd53feb · outbound

This paper cites Maskgit: Masked generative image transformer.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Maskgit: Masked generative image transformer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:31.680617Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:31.680617Z digest=sha256:91900d038c7f3c52e24592a888bd9a26e7de06848688dc97149e1638930b420f

Observation ab85a340-a5c4-4404-a48d-c2a01f7d38e5 · outbound

This paper cites Muse: Text-To-Image Generation via Masked Generative Transformers.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Muse: Text-To-Image Generation via Masked Generative Transformers

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:31.725029Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:31.725029Z digest=sha256:1869ede954bd526d91b47b8448d561372a0aa02d9cb75210e6b78d3f27eea4e2

Observation cb519f03-a043-4030-8e65-b652ae1f883a · outbound

This paper cites ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models ShareGPT-4o-Image: Aligning Multimodal Models with GPT-4o-Level Image Generation

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:31.813465Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:31.813465Z digest=sha256:12597cba957ce22baf30b0cc091eecf2545b0d596ebed78da0446877f5a97642

Observation cdba4620-f658-4390-a193-ec99267fe35a · outbound

This paper cites BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models BLIP3-o: A Family of Fully Open Unified Multimodal Models-Architecture, Training and Dataset

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:31.888731Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:31.888731Z digest=sha256:eb776a113f4d96315929eca524201555222f33c0a992e8df3aa8c5b02c7b6c87

Observation a27ce837-6cb6-4020-9de9-132ab1b3ab10 · outbound

This paper cites Sdar: A synergistic diffusion-autoregression paradigm for scalable sequence generation.arXiv preprint arXiv:2510.06303, 2025.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Sdar: A synergistic diffusion-autoregression paradigm for scalable sequence generation.arXiv preprint arXiv:2510.06303, 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:31.960189Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:31.960189Z digest=sha256:a472114140d7c1527c4e00feff754086572cc1638be99dff5144b5e8550440f5

Observation aa860c1c-2632-4aa7-b6de-71350afa582d · outbound

This paper cites Emerging Properties in Unified Multimodal Pretraining.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Emerging Properties in Unified Multimodal Pretraining

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:32.006411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:32.006411Z digest=sha256:ceef8354ee6a65203f61e634fdb2841f5bccd730807cc6cd9d2e8f11d9a08174

Observation 517b9884-a8b9-4233-b9b7-e4f826f1e19a · outbound

This paper cites Bert: Pre-training of deep bidirectional trans- formers for language understanding.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Bert: Pre-training of deep bidirectional trans- formers for language understanding

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:32.108651Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:32.108651Z digest=sha256:0328a14cf1b64da91cba30366b0de9dd70bd4eceb21dd8730a5538a4dcd2b3f5

Observation 5127ee4f-d1df-4b35-920f-b4421a01b24f · outbound

This paper cites Taming transformers for high-resolution image synthesis.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Taming transformers for high-resolution image synthesis

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:32.180195Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:32.180195Z digest=sha256:dffb250587e77023d0e6d5aadb2a28b46767ef6127b6458ddc0a1d1ee043030e

Observation 4ed01278-b7bd-46de-a1b9-3f8798c0af15 · outbound

This paper cites Scaling recti- fied flow transformers for high-resolution image synthesis.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Scaling recti- fied flow transformers for high-resolution image synthesis

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:32.248872Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:32.248872Z digest=sha256:9b2e7d85e5aacda9c9cede72056b37538e5aa0cd6f3454a997337026fc8eed4f

Observation 00d7975a-5547-4546-9ac3-d260e33ce698 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:32.331138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:32.331138Z digest=sha256:10fe24d03baec7740412022c34e681efa67c1203a0de3560ac9888250d0f349a

Observation fcb18a22-f58e-4f29-bfa1-5ba7f8de4c4e · outbound

This paper cites Geneval: An object-focused framework for evaluating text- to-image alignment.Advances in Neural Information Pro- cessing Systems, 36:52132–52152, 2023.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Geneval: An object-focused framework for evaluating text- to-image alignment.Advances in Neural Information Pro- cessing Systems, 36:52132–52152, 2023

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:32.411609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:32.411609Z digest=sha256:6eddae8a60db8f09f0e3850036db47d4fa71fac608403f44b46e7a21c5fc1a9f

Observation 422dbeb9-5e7f-4fe7-90ca-88f2e6ae6291 · outbound

This paper cites MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models MAmmoTH-VL: Eliciting Multimodal Reasoning with Instruction Tuning at Scale

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:32.491974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:32.491974Z digest=sha256:72ca2094a90cb6a27e43583a45805738b80ee0001bb03254b62f7650a0545fa3

Observation 99356766-f590-4816-a53d-83d6389bb270 · outbound

This paper cites Masked autoencoders are scalable vision learners.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Masked autoencoders are scalable vision learners

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:32.577277Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:32.577277Z digest=sha256:469ce757418784e1899b1d36b18fd1d991ef7cb40ec307891358611224b8add7

Observation 1f0b3f9f-1175-4df1-b673-01756b972405 · outbound

This paper cites Unified discrete diffusion for si- multaneous vision-language generation.arXiv, 2022.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Unified discrete diffusion for si- multaneous vision-language generation.arXiv, 2022

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:32.682171Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:32.682171Z digest=sha256:df79e60b6f7e08762e101779a8caf04049a2cd333da1315f3586457fd49b520e

Observation d869eb79-89e4-4c0a-8a4f-d5441ec53756 · outbound

This paper cites ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models ELLA: Equip Diffusion Models with LLM for Enhanced Semantic Alignment

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:32.763858Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:32.763858Z digest=sha256:a83c5122c305e41c3f2db3dab8e154df51c14b722a99acd5b92b0d20f7a5b414

Observation 50d1afc4-20bf-4fc3-9f59-3eb70eee4f0f · outbound

This paper cites VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models VisualWebInstruct: Scaling up Multimodal Instruction Data through Web Search

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:32.823096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:32.823096Z digest=sha256:0b99a9114eec0d81cfcc07132acacb7156307baa12e24c09c2d6eb59f6f2ab8d

Observation 7ce2fe76-20f9-45e5-b0cb-6fd6b23e069f · outbound

This paper cites Referitgame: Referring to objects in pho- 9 tographs of natural scenes.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Referitgame: Referring to objects in pho- 9 tographs of natural scenes

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:32.907400Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:32.907400Z digest=sha256:b928d83c7bd879eb9f61f1f12c7aaac8bd4fbc71a108bfbe35f6c0e7d33a65b8

Observation e401cbdf-a595-4065-a582-5d2a856263b2 · outbound

This paper cites Mercury: Ultra-Fast Language Models Based on Diffusion.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Mercury: Ultra-Fast Language Models Based on Diffusion

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:33.027701Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:33.027701Z digest=sha256:f592ee1fc0a5119481d9c0b9e866014cbdd9aef92a66981489f318d30d59def7

Observation 77332f7e-e738-4752-b78c-98145d5da34a · outbound

This paper cites Segment any- thing.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Segment any- thing

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:33.150937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:33.150937Z digest=sha256:9f001f3df00a0707cf21e6e77395325b24853ebf20f580ae24a073816992af53

Observation 319a7461-f406-49ff-adee-f0d019eaeea4 · outbound

This paper cites Pick-a-pic: An open dataset of user preferences for text-to-image generation.Ad- vances in neural information processing systems, 36:36652– 36663, 2023.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Pick-a-pic: An open dataset of user preferences for text-to-image generation.Ad- vances in neural information processing systems, 36:36652– 36663, 2023

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:33.359001Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:33.359001Z digest=sha256:9c32ec02d47d778ff445277f134a0901e1cf515082a6c628f457bc13448b72c9

Observation d5b1fb5f-ea04-4549-8986-6a3292b41be0 · outbound

This paper cites Flux.https://github.com/ black-forest-labs/flux, 2024.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Flux.https://github.com/ black-forest-labs/flux, 2024

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:33.430680Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:33.430680Z digest=sha256:d22eaf1e70a785d397c081742b7ee2cdb335fec3635a8a2e65c70175ef267e2d

Observation 73c32410-1435-4f50-9680-efa9adad3911 · outbound

This paper cites Flux.1 kontext: Flow matching for in-context image generation and editing in latent space,.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Flux.1 kontext: Flow matching for in-context image generation and editing in latent space,

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:33.545236Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:33.545236Z digest=sha256:36697741d8b08ef2b14f9a0a4325089d5e7cf5cadb0893e2c373b7f5ec94b671

Observation 690b09b8-36da-40f3-83fb-250bb252f32e · outbound

This paper cites Blip-diffusion: Pre- trained subject representation for controllable text-to-image generation and editing.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Blip-diffusion: Pre- trained subject representation for controllable text-to-image generation and editing

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:33.658393Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:33.658393Z digest=sha256:b1bfc8d40daed76aba472e276efaca826f2b6e45a684bc8e8774b1b00e8b41dd

Observation 0729ccc2-cdc1-4158-b201-4cc182163957 · outbound

This paper cites Playground v2.5: Three insights towards enhancing aesthetic quality in text-to-image genera- tion, 2024.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Playground v2.5: Three insights towards enhancing aesthetic quality in text-to-image genera- tion, 2024

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:33.772773Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:33.772773Z digest=sha256:9235693a9ad9e6d92d59b33639bb68ebd89e003189e3e0b820429ce4ee4ac424

Observation c627d780-daa7-40c1-a254-bbb7bda83ec6 · outbound

This paper cites InstructAny2Pix: Flexible Visual Editing via Multimodal Instruction Following.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models InstructAny2Pix: Flexible Visual Editing via Multimodal Instruction Following

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:33.900776Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:33.900776Z digest=sha256:19c8501f98b44d57bb9f3e6e628f0815465636eb961d4625c92e143060dd9723

Observation 9ba97daa-8a33-4978-838c-beda1e91ac46 · outbound

This paper cites Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Lavida-O: Elastic Large Masked Diffusion Models for Unified Multimodal Understanding and Generation

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:34.063352Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:34.063352Z digest=sha256:19a0f11fbd9df7ff3872f25c65e93f2124374bbc21c0a28e74be7f4d7c16ce08

Observation 7cbfe9af-adca-4c65-a0d1-903bc5d42745 · outbound

This paper cites LaViDa: A Large Diffusion Language Model for Multimodal Understanding.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models LaViDa: A Large Diffusion Language Model for Multimodal Understanding

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:34.225411Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:34.225411Z digest=sha256:922f81a25ff4a1bfe9a22fbf016058fa3c96b18c3953f95596a1e9684ab49e40

Observation aaca2245-3af9-4222-a213-d487d9f7ac8a · outbound

This paper cites UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models UniWorld-V1: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:34.439620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:34.439620Z digest=sha256:2317d4a7fb00922b9f354955a9602233c738e4ef520f2aeaa9ba039b5193efba

Observation 2daaf990-efe9-470a-b0ed-5c5c4adae800 · outbound

This paper cites Grounding dino: Marrying dino with grounded pre-training for open-set object detection.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Grounding dino: Marrying dino with grounded pre-training for open-set object detection

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:34.515666Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:34.515666Z digest=sha256:fc548da75ce004876aad668cd1d60085ce437a82c653b813653c4453889f8da5

Observation 7e34dbe6-471a-411c-9ff4-91460f2a828c · outbound

This paper cites Step1X-Edit: A Practical Framework for General Image Editing.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Step1X-Edit: A Practical Framework for General Image Editing

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:34.652312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:34.652312Z digest=sha256:4b45157ee817170173e975ba98603c6a94ba14cef168db6fc3d49f1eb99e846a

Observation d471d982-2de5-4bd9-9b53-339983219a36 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vi- sion, pages 216–233.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vi- sion, pages 216–233

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:34.806510Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:34.806510Z digest=sha256:f4967bf7befc8a51d3359684398160498213f7f96a8dbb90090b0ba966658b78

Observation e58941af-e7d3-4b9a-bd11-aaa7f1a7b367 · outbound

This paper cites Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Discrete Diffusion Modeling by Estimating the Ratios of the Data Distribution

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:34.918576Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:34.918576Z digest=sha256:e8ee1f52515e5a713a822f57378432a24123e9b657a69c91e68b58cbf1db88c0

Observation b6904efb-a30f-47e1-b0fe-fb6c0a7245f3 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:35.058138Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:35.058138Z digest=sha256:39db782849818504ed3573cc06fbcb04f2d942e411f8805c7bdeca81816e5750

Observation 98901624-b232-4ac6-bf9e-6110fd73e9d1 · outbound

This paper cites Unitok: a unified tokenizer for visual generation and understanding.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Unitok: a unified tokenizer for visual generation and understanding

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:35.155487Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:35.155487Z digest=sha256:1646148060896a8e79a4ae7f3add9d282ad041662eaadff0fb4f37552f9d98e8

Observation c1cad497-ab5f-4dac-90b2-9ee8ba46a4b2 · outbound

This paper cites dKV-Cache: The Cache for Diffusion Language Models.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models dKV-Cache: The Cache for Diffusion Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:35.294723Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:35.294723Z digest=sha256:9f77ffdc540b7ca2e05fd961b941ea43b66b1d260757427a4903545ee66a0236

Observation ec52f14f-8485-46f8-a7eb-ff8ca43e6280 · outbound

This paper cites Hpsv3: Towards wide-spectrum human preference score.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Hpsv3: Towards wide-spectrum human preference score

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:35.412047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:35.412047Z digest=sha256:351fa90138164c7b7643ebd4e934bfe85cb2993fa992d7d453f862c8cbcbd3f2

Observation 63ee8488-85c0-405f-8172-dc7d1c1e489d · outbound

This paper cites ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models ChartQA: A benchmark for question answer- ing about charts with visual and logical reasoning

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:35.568594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:35.568594Z digest=sha256:01a0f416c6d45d054d25296493f560596292d5e0fe5cf0ef6136586b838fe60e

Observation 03b93bcc-167a-4913-8875-eff802353be3 · outbound

This paper cites Docvqa: A dataset for vqa on document images.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Docvqa: A dataset for vqa on document images

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:35.722711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:35.722711Z digest=sha256:73bd4f96a5579faf1fd1e6bbc5b9a7b2f698d00e3d901c8b22c84e8f53e3d684

Observation e2f6a1be-aeb2-4557-95ef-cc8e61c54de9 · outbound

This paper cites Large Language Diffusion Models.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Large Language Diffusion Models

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:35.889523Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:35.889523Z digest=sha256:06b2459e4137ab01e6c834bb6f9abe5014f812ea69c19275d6f3bdd97850766c

Observation 63bc66d1-e3c9-4aca-952a-516a8e378b55 · outbound

This paper cites Dall·e 3.https://openai.com/index/ dall-e-3/, 2023.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Dall·e 3.https://openai.com/index/ dall-e-3/, 2023

Reference 46

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:36.053626Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:36.053626Z digest=sha256:661df51ac7612022523868f5674ef2ab63aea6fb7834ebdb9bdf45c3ec7bebef

Observation e1f20f59-01b6-44d5-8f6a-8217356ed24b · outbound

This paper cites GPT-4o System Card.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models GPT-4o System Card

Reference 47

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:36.170362Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:36.170362Z digest=sha256:ae6a13c7bdbe6c23487f29c5d68f01bb018349a12c5a18804cb5b4caafe17320

Observation 00c49ad6-72fb-4335-bb25-d856fabe1332 · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:36.303981Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:36.303981Z digest=sha256:d0baa9c5f135bc261d1dd520e1fb49e3a3782225ca548b18adc3bfdb59efd33f

Observation f94b5626-92a6-4298-8f75-95767ead62d9 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Learning transferable visual models from natural language supervi- sion

Reference 49

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:36.424148Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:36.424148Z digest=sha256:849828f7d90fca58eab4a3f9b83d827d3d54e52d2344b6f53a29ce2b6982285f

Observation 15eb1e5e-b42f-4a28-8b28-91817e75bc80 · outbound

This paper cites Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:36.537363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:36.537363Z digest=sha256:cc6c681bf05b91691e8588722568fa13e5a207d59837e4021f027509794dce48

Observation f32406eb-1785-4d3d-a436-dc26e5c8f9b0 · outbound

This paper cites Simple and effective masked dif- fusion language models.Advances in Neural Information Processing Systems, 37:130136–130184, 2024.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Simple and effective masked dif- fusion language models.Advances in Neural Information Processing Systems, 37:130136–130184, 2024

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:36.672279Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:36.672279Z digest=sha256:0fe654b8907068aef8535554a5a9678c4bcfc9048436fdedaf34fa9790361263

Observation 592a8b64-1010-4728-8184-8c1c3dc37b68 · outbound

This paper cites Laion-aesthetics.https : / / laion.ai/blog/laion- aesthetics/, 2022.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Laion-aesthetics.https : / / laion.ai/blog/laion- aesthetics/, 2022

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:36.822838Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:36.822838Z digest=sha256:8fcb4b374a9acfa1bbcf54ea6b000c068499165a6ac1923ce4e190138aa014a9

Observation 260c8988-8800-44b3-8b33-64017bb8485b · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural in- formation processing systems, 35:25278–25294, 2022.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Laion-5b: An open large-scale dataset for training next generation image-text models.Advances in neural in- formation processing systems, 35:25278–25294, 2022

Reference 53

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:36.935037Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:36.935037Z digest=sha256:ed225fdba39999de2a967dfda453102ed9cf2fc788f96bc686ebe731307532f3

Observation 8580dfb8-ab89-4628-8b9d-f54d23b16015 · outbound

This paper cites Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Muddit: Liberating Generation Beyond Text-to-Image with a Unified Discrete Diffusion Model

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:37.055379Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:37.055379Z digest=sha256:de5d4bea86ebaf3688c02378ae6186e5e25edae9865c2c125fcad44070f882e8

Observation 422723de-f80c-4a2a-9eec-9d74d89d962c · outbound

This paper cites Sparse-dllm: Accelerating diffusion llms with dynamic cache eviction.arXiv preprint arXiv:2508.02558, 2025.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Sparse-dllm: Accelerating diffusion llms with dynamic cache eviction.arXiv preprint arXiv:2508.02558, 2025

Reference 55

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:37.195079Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:37.195079Z digest=sha256:d26454f6c9758e6be2d9a086ec3fa4a2b13f4479d426716f34f413cdd9355baa

Observation 10be79d1-d1ac-43bc-9579-4646383e9453 · outbound

This paper cites Journeydb: A benchmark for generative im- age understanding.Advances in neural information process- ing systems, 36:49659–49678, 2023.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Journeydb: A benchmark for generative im- age understanding.Advances in neural information process- ing systems, 36:49659–49678, 2023

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:37.313959Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:37.313959Z digest=sha256:1da94fdbad69b9cafae13a0d98b92f3fd5a763a33f6fe9da84ddd604194b8ec5

Observation 4febda56-29fd-4f78-bf8d-e93b4ac5d848 · outbound

This paper cites Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Diffusion LLMs Can Do Faster-Than-AR Inference via Discrete Diffusion Forcing

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:37.455637Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:37.455637Z digest=sha256:299657c3c115310f416673b0f8908657f3712e362f289749a6d30805cda618c4

Observation 159cbec4-d1ea-40e7-8c4e-790c2f00a3e8 · outbound

This paper cites Segllm: Multi-round reasoning segmentation with large language models.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Segllm: Multi-round reasoning segmentation with large language models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:37.549253Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:37.549253Z digest=sha256:8f5dbda4bebca323fc6daf19594096acaca2c961d7dc4ecdb3645bc961c1a51b

Observation f30cd611-26b5-44df-9794-0973834a8c51 · outbound

This paper cites GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models GPT-IMAGE-EDIT-1.5M: A Million-Scale, GPT-Generated Image Dataset

Reference 59

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:37.664980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:37.664980Z digest=sha256:d3a19e221071a87e5c95bf45553774faa0b26fcae5ce7089191d1ea131600b26

Observation a3ae9a3f-10e6-47fc-9d0f-9c9a8e73f11d · outbound

This paper cites Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Fast-dLLM: Training-free Acceleration of Diffusion LLM by Enabling KV Cache and Parallel Decoding

Reference 60

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:37.768727Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:37.768727Z digest=sha256:9213b13c17bf474d6ff32e107eb37d4d845982d6a297a6bedc981fb5e03b6cd0

Observation 9fa80018-870b-444a-9a83-36ae5772908a · outbound

This paper cites OmniGen2: Towards Instruction-Aligned Multimodal Generation.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models OmniGen2: Towards Instruction-Aligned Multimodal Generation

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:37.852349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:37.852349Z digest=sha256:f6f0d2cfc372c97adb0b4966870ab363af90c2693fcac12b08f0175b0e724efa

Observation 381f95c8-fbb1-4bfb-883b-73b41cd9acdb · outbound

This paper cites Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Human Preference Score v2: A Solid Benchmark for Evaluating Human Preferences of Text-to-Image Synthesis

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:37.920408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:37.920408Z digest=sha256:a3a910b802929c4273b05e16db8454c5fbe972e06df91849586085b494661065

Observation 7bb4e548-69f4-4d12-885c-86d4bc19d753 · outbound

This paper cites VILA-u: a unified foun- dation model integrating visual understanding and genera- tion.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models VILA-u: a unified foun- dation model integrating visual understanding and genera- tion

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:37.984425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:37.984425Z digest=sha256:6cc466aa6e858c47b3fc26d5dcadec0e9549448c1c8034941555fc378ce89226

Observation f1396d73-a551-442e-ba77-8054db4ddb14 · outbound

This paper cites Omnigen: Unified image genera- tion.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Omnigen: Unified image genera- tion

Reference 64

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:38.047544Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:38.047544Z digest=sha256:d7b3de37dc662934c7d920b6be097596cb225185931364e5c83e35cb179136f3

Observation 722332ec-76dd-4aec-b09e-56fc90c16713 · outbound

This paper cites Show-o: One Single Transformer to Unify Multimodal Understanding and Generation.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Show-o: One Single Transformer to Unify Multimodal Understanding and Generation

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:38.152198Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:38.152198Z digest=sha256:e606bb83bc89030b8ac48fa4fe7574f4b2b34fdff018157fa76be25046c9af04

Observation 1dd19a22-ab81-4af8-b0e9-42b4d1a2685c · outbound

This paper cites MMaDA: Multimodal Large Diffusion Language Models.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models MMaDA: Multimodal Large Diffusion Language Models

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:38.232907Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:38.232907Z digest=sha256:559b093b32dfdfc8a89683f0d6a2d0c9f01e45ef3425ef1f31b4750d77f10504

Observation ac717bf6-846d-4940-bdaa-942dc6dc7261 · outbound

This paper cites Dream 7b,.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Dream 7b,

Reference 67

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:38.338674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:38.338674Z digest=sha256:c65834e2b21efa6cd029e575f6446c65d702284c2815183aebb35c629418e0f9

Observation 0e12e5bf-806c-4ee4-8530-c61d9e04d8cb · outbound

This paper cites ImgEdit: A Unified Image Editing Dataset and Benchmark.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models ImgEdit: A Unified Image Editing Dataset and Benchmark

Reference 68

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:38.419317Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:38.419317Z digest=sha256:c88a207eaefe9b8fbbbff6e7cc86e0401f8ac059a901a0bd97e6fc493289b610

Observation de50c696-4650-4c37-a92b-b230fba37457 · outbound

This paper cites LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 69

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:38.516268Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:38.516268Z digest=sha256:d882bef666d642b1b741d687b888c0c738834bd1d07be8f23ade42a9a9cb278a

Observation 2ed0f6df-0948-4e1f-8440-92a1dcc99d6b · outbound

This paper cites Anyedit: Mastering unified high-quality image editing for any idea.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Anyedit: Mastering unified high-quality image editing for any idea

Reference 70

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:38.637196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:38.637196Z digest=sha256:8399734f03f092f4ffe3810159f0e98e8d8cf0cbb287ced82216a99802b6134c

Observation ca2e4a40-27bc-4821-8a89-c6dca9e1d453 · outbound

This paper cites Dimple: Discrete Diffusion Multimodal Large Language Model with Parallel Decoding.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Dimple: Discrete Diffusion Multimodal Large Language Model with Parallel Decoding

Reference 71

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:38.748437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:38.748437Z digest=sha256:fe3a6239234c71a5940d12dd44acd6998e4e2712700bbb89ca29867848d10dd6

Observation 54202b41-ca23-432e-8b6d-8dd4638d97b2 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi

Reference 72

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:38.857913Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:38.857913Z digest=sha256:01990dae38f87996bf82b9229c252f182fe2f5538b0fde087cd494131e2dd7e0

Observation c909dbad-c9b2-490e-a03c-53a47f7ce459 · outbound

This paper cites Magicbrush: A manually annotated dataset for instruction- 11 guided image editing.Advances in Neural Information Pro- cessing Systems, 36:31428–31449, 2023.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Magicbrush: A manually annotated dataset for instruction- 11 guided image editing.Advances in Neural Information Pro- cessing Systems, 36:31428–31449, 2023

Reference 73

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:39.024404Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:39.024404Z digest=sha256:288c9c80e5e0a3d0b7b1796006549acf30d89b56a591d303ca82174f6441c8c8

Observation 9f36e4c7-411e-4edb-a80f-fdd43ca7f9d8 · outbound

This paper cites MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models MathVerse: Does Your Multi-modal LLM Truly See the Diagrams in Visual Math Problems?

Reference 74

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:39.116064Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:39.116064Z digest=sha256:2900e34a4437f2012ef37619160d27bd5731af39b11e4375d1ae4458a5d5b16c

Observation 007b409c-f1b0-444b-ad40-dc6f57e7ab02 · outbound

This paper cites Ultraedit: Instruction-based fine-grained image editing at scale.Advances in Neural Information Pro- cessing Systems, 37:3058–3093, 2024.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Ultraedit: Instruction-based fine-grained image editing at scale.Advances in Neural Information Pro- cessing Systems, 37:3058–3093, 2024

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:39.209645Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:39.209645Z digest=sha256:e7d66f2d28cc0ff87052b1e3778d442c57806e05b26afbd6f291726850d7ae5b

Observation 8e9ed9ad-2b51-4d19-9f3a-b206f545faad · outbound

This paper cites Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model

Reference 76

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:39.318402Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:39.318402Z digest=sha256:e28e296503fb4eab5dee26ef8447ffd5def459100250818806f2da523478219e

Observation b5ff52f0-cf5f-4df5-afba-fa5eba706c45 · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 77

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:39.471015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:39.471015Z digest=sha256:0f1f24fae002ed905da00958a0e9a8ef6dfec832c00a6c5d771bbdf46eeb3da7

Observation 5d2662de-ecbb-4b08-b19b-874cbf05d83e · outbound

This paper cites an unresolved cited work.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Unresolved cited work

Reference 78

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:39.643580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:39.643580Z digest=sha256:8cdfdf245e53a19965cdc3de42ccbc010056a35141e36683d33de21c84b37a93

Observation b871aece-7911-4e31-98a6-59f15ba397d9 · outbound

This paper cites an unresolved cited work.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Unresolved cited work

Reference 79

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:39.884046Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:39.884046Z digest=sha256:3ec62f0e231f80a0ea1e7103e3188b0ed2812774fb02b0e73a9dfd07688075e1

Observation 5e4aec08-24b3-4f5e-b7b7-466f99a29243 · outbound

This paper cites [reg]” that is similar to mask token “[M].

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models [reg]” that is similar to mask token “[M]

Reference 80

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:40.089008Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:40.089008Z digest=sha256:33a4d5376ef8cda7ab230feb614d7759219a949dfe6ea6ee82c25818a7a675db

Observation 695b6728-2c0b-4029-977a-41e9968af7bb · outbound

This paper cites Data pipeline Our training dataset consist of the following tasks.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models Data pipeline Our training dataset consist of the following tasks

Reference 81

Resolution
malformed identifier
no resolver link, observed 2026-08-03T16:21:40.175124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:40.175124Z digest=sha256:f73cfd144aea9b100ededf606c4ddd83d97947449fb5e44133363b9f64bb97d9

Observation 4258395d-0a80-4fe5-aaa0-daa6118d9e71 · outbound

This paper cites First, the speedup only benefits long sequence generation such as text-to-image generation, image-editing, or visual math problem-solving with long reasoning chains.

Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models First, the speedup only benefits long sequence generation such as text-to-image generation, image-editing, or visual math problem-solving with long reasoning chains

Reference 82

Resolution
unresolved
no resolver link, observed 2026-08-03T16:21:40.277901Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T16:21:40.277901Z digest=sha256:be11c5480aad587b03591feb029ee5faaef5bacae5a0fc252138edd3c09765f3

Pith citing papers

Observation 7629bb48-de80-405d-94d8-eac60c43604b · inbound

Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models cites this paper.

Visual-Redundancy-Controlled Parallel Decoding for Diffusion-Based Multimodal Large Language Models Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-17T01:20:41.788401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-29T23:09:32.594194Z digest=sha256:2124204cfcdce5d19c07cd5320075f0f63bfbdaaf852826c8a664eb018c9eab6

Observation af62d3ef-f720-4a86-9729-568b77fdb0a7 · inbound

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis cites this paper.

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-07-17T01:20:41.788401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=pdf_text observed=2026-06-30T06:07:55.600338Z digest=sha256:37c51cdabef028f3a0d84fe2b4f7fda781b3ac7f8384bdd190de552b06752a34

Observation 984d7d8e-7210-4a8e-8593-09f46d87463b · inbound

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis cites this paper.

Nemotron-Labs-Diffusion-Image: Advancing Masked Discrete Diffusion for High-Resolution Image Synthesis Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-02T09:39:33.093629Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T09:39:33.093629Z digest=sha256:d53fa7a47385f91c527956ceaed012d4ef3a385f0bb19a2c46f43437cf45ba3e

Observation 1e3fa0bf-8eb9-4b00-870e-e304e1d71f41 · inbound

Set Diffusion: Interpolating Token Orderings Between Autoregression and Diffusion for Fast and Flexible Decoding cites this paper.

Set Diffusion: Interpolating Token Orderings Between Autoregression and Diffusion for Fast and Flexible Decoding Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models

Reference 134

Resolution
verified exact
arxiv_id, observed 2026-07-17T01:20:41.788401Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-14T06:32:32.682623+00:00.

source=arxiv_source observed=2026-07-03T17:30:39.458521Z digest=sha256:5d95673d9250e48d7b177bafcbce4061f22fe01ea8d28a6ca03f34fcd3fdce79

Observation cc4f3ba4-420e-41ed-999f-08fee99c6af5 · inbound

Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation cites this paper.

Seeing the End at Step Zero: Accelerating Diffusion MLLMs via MLP Sparsity-Aware Truncation Sparse-LaViDa: Sparse Multimodal Discrete Diffusion Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-02T01:48:51.932358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:48:51.932358Z digest=sha256:4709710a0d35bd0138054f2a409e6b218002c2a4e009bb7658d4f6f408cc5f67