Pith. sign in

Paper Citation Record · LEDGER

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models

As of 7 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2606.19534.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2606.19534 v1

Coverage vector

measured 62 of 62 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-06-26T20:59:26.886235Z

measured 63 of 63 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-07-14T04:38:05.237334Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

62 of 62 outbound references displayed

  • verified exact27
  • verified fuzzy0
  • unresolved33
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch2

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 69465d34-1e91-4d0f-8be4-fcbe807e7c05 · outbound

This paper cites LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training

Reference 1

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:17.966899Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:35c34bb92763f57507fd8c28b297518337db8662dcfd60c607ead9a9006662d7

Observation edcd7512-d4b4-45b2-9d2a-0778e3952b64 · outbound

This paper cites Qwen3-VL Technical Report.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Qwen3-VL Technical Report

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:18.034990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:a03aa9fb70b8ef7b91053eac25983b14b80fa90931c1d9226def29a5dd536969

Observation c6f416eb-220a-4221-86e9-a94bf855832b · outbound

This paper cites Qwen2.5-VL Technical Report.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Qwen2.5-VL Technical Report

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:18.030383Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:808068b5a5d91f5376d302428d5208fc7d8fed84e57a98e460db80257b24bdd9

Observation dc8738d5-5022-437b-ac97-d5cc843b92b1 · outbound

This paper cites LLaDA2.0: Scaling Up Diffusion Language Models to 100B.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models LLaDA2.0: Scaling Up Diffusion Language Models to 100B

Reference 4

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:18.039142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:2f0f015e9adfdd9fc244762de4e0c5d4709b5a5910a2ff7d8fbbfe91548c08eb

Observation 60a76bbc-69c0-44e6-ab0f-0fcb8e0ae680 · outbound

This paper cites SAM 3: Segment Anything with Concepts.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models SAM 3: Segment Anything with Concepts

Reference 5

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:17.957492Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:73d7670bb444a2850c6db2baef0adc64ef2e1cad087117a4549c7bb83ba03e01

Observation f074caa9-2d56-40fd-9007-920e4ba0fb36 · outbound

This paper cites Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic

Reference 6

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:18.042882Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:10bcbf7762e8bc0a1715077839daa97a9fe779acfbb34dd623dab4d515b3a3bd

Observation 81a74e3f-d806-4c02-a06d-66b480a1592c · outbound

This paper cites Are we on the right way for evaluating large vision-language models? InNeurIPS, 2024.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Are we on the right way for evaluating large vision-language models? InNeurIPS, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:b431b53f1a305a5294b2c1cd3a4357a6161dba401a6f5855de6e6672c1ce6103

Observation 32a425e3-c04f-4d7f-b7f3-01c3542cbf20 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:17.941439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:b547a2215f7626860709c64dcd125543ca42877689fba57f8608da181159cdec

Observation 69c4fb90-6787-46bb-8740-8e444176e815 · outbound

This paper cites Sdar: A syn- ergistic diffusion-autoregression paradigm for scalable sequence generation.arXiv preprint arXiv:2510.06303.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Sdar: A syn- ergistic diffusion-autoregression paradigm for scalable sequence generation.arXiv preprint arXiv:2510.06303

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:17.971828Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:eb52d318559f815dd8f0eaa5ac8f12f73cdf95d48ced598c7fda4ab41a556806

Observation 1034c832-6371-4d35-bb79-82a7104ad9f3 · outbound

This paper cites Coconut: Modernizing coco segmentation.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Coconut: Modernizing coco segmentation

Reference 10

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:ed0cfa90f7579836d334c2a260155a6353f19c9c7d2e4505e4e25746e7dd2e16

Observation 1bcbb7fb-6833-411b-bd3e-a0fb4fd980d2 · outbound

This paper cites VLMevalKit: An open-source toolkit for evaluating large multi-modality models.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models VLMevalKit: An open-source toolkit for evaluating large multi-modality models

Reference 11

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:bd991352b3a9fbbd3c18cae541235d103ff4ef6eae7b47071236fcbc88defc09

Observation b7065d67-8233-4124-bfdf-2f89ce39d9a4 · outbound

This paper cites Blink: Multimodal large language models can see but not perceive.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Blink: Multimodal large language models can see but not perceive

Reference 12

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:ec28d0153ee334cf804e423c4f2ba5b798a3597e8b4663cd63f9ce4be09cbef7

Observation 985f5488-3e60-48a4-be6e-710af8944d4d · outbound

This paper cites The Llama 3 Herd of Models.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models The Llama 3 Herd of Models

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:18.027794Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:991a11247cedb0b3b50fe522db587d18c4723a6b9d1badc5cec282a53036b4d4

Observation 6addfe4d-986b-410d-ae39-3528dbdf7ee1 · outbound

This paper cites Dataseg: Taming a universal multi-dataset multi-task segmentation model.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Dataseg: Taming a universal multi-dataset multi-task segmentation model

Reference 14

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:17265bbf5b1c88c8a0e91a789d3d2317a6f8ce27be4ffbefbaed14231f383332

Observation 4092cf1b-cc5e-4ae6-84ae-ec7ac589face · outbound

This paper cites Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:658b6df65ee3897d48ce58a6bfb6dcb2040ddf3765c53104168b314c7f3d0765

Observation 8ea77c65-c133-45dc-a0e2-372d83264c7d · outbound

This paper cites DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning

Reference 16

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:18.054172Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:2900069606faafa3022a4672d8f384cd03a93c90011b09470df392bf9aab3ba7

Observation c9ff9b25-681c-4432-8761-78883a39ac0d · outbound

This paper cites Ai2d: A dataset for diagram understanding.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Ai2d: A dataset for diagram understanding

Reference 17

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:5a00329d1b753171a157bd2c007dd4c6ce759add0f765ce4fead78c1ca0f39b3

Observation ed7700ed-c843-4e75-8f84-02553ee9164b · outbound

This paper cites Segment anything.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Segment anything

Reference 18

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:e0e9762d11b1afd91ac70e7bfa6f0746691177f62907469d7f005effc0b0bb9c

Observation 7c1e478e-961e-4009-9c0d-72a8061d46e4 · outbound

This paper cites The scalability of simplicity: Empirical analysis of vision-language learning with a single transformer.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models The scalability of simplicity: Empirical analysis of vision-language learning with a single transformer

Reference 19

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:9ba7525854f75383229e1d41818e93e242a32f8d1bc652df3164d29cd7a3b467

Observation caf6e1e2-57b0-41ee-a72e-d0e331a3545a · outbound

This paper cites Seed-bench: Benchmarking multimodal large language models.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Seed-bench: Benchmarking multimodal large language models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:1d813ff5310a75a219ca3d6f3b6248e3af606dd9baffd8512e772765d80ff883

Observation 7f702793-1820-4b27-9119-09987a479169 · outbound

This paper cites Llava-onevision: Easy visual task transfer.TMLR, 2025.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Llava-onevision: Easy visual task transfer.TMLR, 2025

Reference 21

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:c5da6b6b87f398e0e9da4459d55e1cf6a87499a0bfc3a6d6a8e4fdff2b289bee

Observation 6d5d9761-0f57-4a1e-915f-e3d09818a77e · outbound

This paper cites DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:17.978718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:fafef6648d259070054ad6da9d513242bc7b6a363da71c81ca438d08cab4202e

Observation eabc9dd0-9e18-4a5b-8158-9306d1033461 · outbound

This paper cites Describe anything: Detailed localized image and video captioning.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Describe anything: Detailed localized image and video captioning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:9d3ffb2c358e61f71ff5d495af9f7968b926effe82f0485083e18b8ace6d56be

Observation d98e4add-d5db-43f3-ab25-04f496bbff13 · outbound

This paper cites Microsoft coco: Common objects in context.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Microsoft coco: Common objects in context

Reference 24

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:ac27feffb0fac80eda405aa7e003d2191fd2436f44197f70af6068a97b83cdfc

Observation bf609126-ac93-4c8d-a453-f751e9d57b4e · outbound

This paper cites Visual instruction tuning.NeurIPS, 2023.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Visual instruction tuning.NeurIPS, 2023

Reference 25

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:f2c427dd66ba3b3f6b9bb5acd1044f816bb0b39cfda9da7a8ba4952d814cb782

Observation 57699af0-374f-489f-8875-5c7dc83de466 · outbound

This paper cites Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pp.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pp

Reference 26

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:f5c0131ffcdc3cccc2f0d85c9a6deb1dcb954c3a30e75df626d22101db715d37

Observation 27bb208d-a5a3-46eb-b473-0b226e5eb6ff · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 27

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:524e5caace669c51d4298c125fec65a18c9fa3916135996774b6e056e0320be0

Observation 3eec8668-8fcb-4299-adf3-82246b1b453e · outbound

This paper cites Chartqa: A benchmark for question answering about charts with visual and logical reasoning.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Chartqa: A benchmark for question answering about charts with visual and logical reasoning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:5f499a53ca2f30f0bdf429a0d27f363d36dcf5167a905ee01d209be4a6f48f86

Observation 75201859-afcf-43f4-83c6-887f0feccea7 · outbound

This paper cites an unresolved cited work.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Unresolved cited work

Reference 29

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:0658bade560a739ad8a700ce575fe698338b0b02e8f696ccdce50bf78cbbde37

Observation 5a9c32be-0c02-4617-93f7-9729fd6f1cfb · outbound

This paper cites InfographicVQA.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models InfographicVQA

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:18.049881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:52f71d65bea95b30a9a64e833c322a66e12fea7ccca34659cdb28670063acc5c

Observation 9fced0f7-1ff4-4aa8-a8ac-d5d7cf8e1356 · outbound

This paper cites Open-o3 video: Grounded video reasoning with explicit spatio-temporal evidence.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Open-o3 video: Grounded video reasoning with explicit spatio-temporal evidence

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:18.041246Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:2b2b515cba1a139257545e6960f57a04098c99f1b8159bbd0bbeb89a41ba8fe6

Observation bb3416f3-4bd8-4940-9bf2-10cd1ddf524f · outbound

This paper cites The flexibility trap: Why arbitrary order limits reasoning potential in diffusion language models.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models The flexibility trap: Why arbitrary order limits reasoning potential in diffusion language models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:c8f379a2c127c453af7d20958a4d7e530de90d559f4b258627cd6aeb075c8a43

Observation 1a5f57a9-298c-47bc-9f48-08a33bf5db61 · outbound

This paper cites Large language diffusion models.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Large language diffusion models

Reference 33

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:8842d050c365340331e80debe2f4c91ee6ced87f26c7e39071e9b356079c6cfe

Observation 365bd1b4-dc9e-4c97-9c5d-4b864b82fe2b · outbound

This paper cites Openai-gpt-5.2.https://openai.com/index/introducing-gpt-5-2/, 2025.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Openai-gpt-5.2.https://openai.com/index/introducing-gpt-5-2/, 2025

Reference 34

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:db713996626c82823afa181bdb259ecf53a274513ef49fda7898e625ff8ada67

Observation 961d46b0-bd76-4b90-bc32-731911991abf · outbound

This paper cites d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-08-07T02:38:17.181363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:0a4d5e25c74f5e30473b4a01cca849ab9530679c63b5f3fca825a6c173d61a09

Observation b69a7643-0ab1-49ae-95e6-8e99ba9cab11 · outbound

This paper cites Glamm: Pixel grounding large multimodal model.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Glamm: Pixel grounding large multimodal model

Reference 36

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:4a17899e78523c6e8c6e6150f5cd20974f1092305d3139d40dac4e87d9b5f026

Observation 40e00f14-7ed4-4f88-a8ec-68372a8c307b · outbound

This paper cites Objects365: A large-scale, high-quality dataset for object detection.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Objects365: A large-scale, high-quality dataset for object detection

Reference 37

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:d1a3e6b460046c8d37ed50f8d387830da4cce50d65914227eed855660d691213

Observation 0262fa8a-2bb8-4cd6-85af-93a6e9077d2d · outbound

This paper cites Cambrian-1: A fully open, vision-centric exploration of multimodal llms.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Cambrian-1: A fully open, vision-centric exploration of multimodal llms

Reference 38

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:2b51bb00cf1da8aa58ce70b2687a1986f624008adbe6f7fa2209483ba3f2257f

Observation 7b3efd9e-afbd-4593-bb19-9b4516a6026e · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 39

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:fb7b09a97c0c89602d3643249c28eef7adf5cc4509b6c0d4b9c78e6bdd7f9717

Observation 599b1598-a3f2-4c7c-9796-fc209c233e0b · outbound

This paper cites SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features

Reference 40

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:17.985767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:75a798891504e162d88fc1a02f43e60fcfaf2090bf9cdd03eaeaa721c268f855

Observation 376d24da-6733-471c-8065-713f086183e2 · outbound

This paper cites Grasp any region: Towards precise, contextual pixel understanding for multimodal llms.ArXiv, abs/2510.18876.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Grasp any region: Towards precise, contextual pixel understanding for multimodal llms.ArXiv, abs/2510.18876

Reference 41

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:18.032446Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:4ad3a7287b667a8925a464bcc31174d87fa44b74ed176a95710da76038eb0f58

Observation 5f12108a-f962-4534-ab2c-fe8eaa15d317 · outbound

This paper cites Ross3d: Reconstructive visual instruction tuning with 3d-awareness.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Ross3d: Reconstructive visual instruction tuning with 3d-awareness

Reference 42

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:449ecc2c5f678d2c45af7d888da530440c12b6e19529775b372253583a39d70b

Observation b0bea613-3961-48fc-b414-37476d1e0217 · outbound

This paper cites V*: Guided visual search as a core mechanism in multimodal llms.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models V*: Guided visual search as a core mechanism in multimodal llms

Reference 43

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:a75fd769c29081425a6c5ab602640e505872d356d53799d97b6a960ac0f3fa99

Observation 3c5f94fe-7cf3-4d5c-a796-4d8d0b1f3f78 · outbound

This paper cites Realworldqa: A benchmark for real-world spatial understanding.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Realworldqa: A benchmark for real-world spatial understanding

Reference 44

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:13536a5a277667595d13aab40c10ac852caf2e6b528353798a8b462f0ec36d1b

Observation 631f3465-7d80-4135-9859-7c214aea7b70 · outbound

This paper cites Lumina-dimoo: An omni diffusion large language model for multi-modal generation and understanding.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Lumina-dimoo: An omni diffusion large language model for multi-modal generation and understanding

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:17.948765Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:a93bbe8bcd7704f5420eda3208faf582cff03650be499a6d2954090040b049ec

Observation 1f1d4d8b-b552-4b90-b73f-f5929da88acc · outbound

This paper cites Qwen3 Technical Report.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Qwen3 Technical Report

Reference 46

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:18.036381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:9ed6c7db433bce13f6506bf9b2e77fa648c62e62be5cfd747abe5aa7bfdcf998

Observation 5526169c-264e-4da9-9212-d3e6b66e7435 · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 47

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:17.976188Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:caa2ba29d68742d5ebf8ac3732e472d82c740f286f37fb398ec11c0a1d0c0399

Observation 9722ad0e-3013-4482-8184-58930a299833 · outbound

This paper cites Mmada: Multimodal large diffusion language models.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Mmada: Multimodal large diffusion language models

Reference 48

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:7b7f00adae4d07dfc55b7bde8802084086be1bf8c4b799e25bb52c20a8dfb5d0

Observation 759bbd20-b2e3-4749-8f45-d7100e30db65 · outbound

This paper cites Dream 7B: Diffusion Large Language Models.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Dream 7B: Diffusion Large Language Models

Reference 49

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:17.955122Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:eee635e89e24082011d2278821b955a470ff2df2056757e93d67ac5b46cae63a

Observation d13cc417-d2c5-4d66-8cfd-cc2f299eca4d · outbound

This paper cites LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning

Reference 50

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:18.047342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:71cd8bee4f72148e2962b08fa89e01765449f5bfbc898a46df47a854fc361496

Observation 7e4f28f1-2df0-4fcd-a6fe-9c177e3daa78 · outbound

This paper cites Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos

Reference 51

Resolution
verified exact
local_arxiv, observed 2026-07-04T00:49:17.981146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:eda5061bec366fd974eaebc31f75d130337976507229420107a2c5755a9d09b2

Observation 2e7bb20e-13d2-4162-98d5-e1c8b9b0587f · outbound

This paper cites Visual reasoning tracer: Object-level grounded reasoning benchmark.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Visual reasoning tracer: Object-level grounded reasoning benchmark

Reference 52

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:18.007437Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:05ae6fa913256c6c10bf50a8a5f1a14e541d94a08d885c4c4a860621b74fd491

Observation 5ced1f86-011a-4c18-a68c-35faa9a26abe · outbound

This paper cites arXiv preprint arXiv:2510.23603 , year=.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models arXiv preprint arXiv:2510.23603 , year=

Reference 53

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:18.045577Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:63284a5622328bb8f86e8ac567df3ac9c0bf26a318eef0abfb814dabc6bdc259

Observation d8e3b033-faad-42a5-9651-3be9e6698292 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 54

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:f9f4553f95c1174d0e1017827191d9eaa4a7ae5a78317703f8f9b7a357c88751

Observation aa64f7c4-f7a0-4cd7-ac01-3ea84e42919b · outbound

This paper cites MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi

Reference 55

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:20c15d6adce00e92b671589f14b74ae9fa910d43cc82f30572b02b2b67967139

Observation db0bbb23-5a06-4e20-a995-698d56181b07 · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In ECCV, 2024.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In ECCV, 2024

Reference 56

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:b223afbef8e1c965ccaecadf5a440cc480971dd327ba40a024ea0cdc42768576

Observation 8965d514-eac4-4041-a629-53856b3c9536 · outbound

This paper cites GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest

Reference 57

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:18.014142Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:998cb571374a5acb7cf186fb076014880e78367867ef063d7c75190aed753392

Observation 371d1551-fea6-4dd8-a715-6250d03aa42a · outbound

This paper cites Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding.NeurIPS, 2024.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding.NeurIPS, 2024

Reference 58

Resolution
unresolved
no resolver link, observed 2026-06-26T20:59:26.886235Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:73bdb2d8644de46abf66d93dbdf8e0d262afbc39e549ecb551ebd6ffade69fe0

Observation cf4aeb30-1799-4ff3-a2a5-7e0bbab87631 · outbound

This paper cites Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding

Reference 59

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T00:49:18.026551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:5463ed6371e4c728b4136c192e3b49f2b7ea689aad5fcf3f79e6ffc0336db768

Observation d4450f81-72ff-4bb2-8715-2f4d3a6ffcdc · outbound

This paper cites Bee: A high-quality corpus and full-stack suite to unlock advanced fully open mllms.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Bee: A high-quality corpus and full-stack suite to unlock advanced fully open mllms

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:18.016324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:89995b30d089014068c952df41847aa7c59c40e055cc71e84edbb4e58fea27f0

Observation 6c463358-a249-4d29-8403-dc4d0dfafc75 · outbound

This paper cites Samtok: Representing any mask with two words.arXiv preprint arXiv:2601.16093, 2026.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Samtok: Representing any mask with two words.arXiv preprint arXiv:2601.16093, 2026

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-07-04T00:49:18.005405Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:21307e1a5310c0599c09e8f10cad026da16058775e47238de5ffaf6f7b626ef7

Observation 7326d3a8-24dd-42ed-9d4a-1d90b0edbc3e · outbound

This paper cites InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models.

PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models

Reference 62

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T00:49:18.003257Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T20:59:26.886235Z digest=sha256:c80ec8fa9e2a40556eace8f2095d3888ca1557699ac270aa81bd90b808edf30c

Pith citing papers

Observation dc62360a-eee7-4d13-8555-e62edb66b550 · inbound

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO cites this paper.

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-07-14T04:38:05.237334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:38:05.237334Z digest=sha256:22d741a5f5db9a06400a30c3e41a83199c0ee7363262fa1566c7f61190846344