Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T20:59:26.886235Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 62 of 62 outbound references and 1 inbound Pith citation observation for arXiv:2606.19534.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-06-26T20:59:26.886235Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-14T04:38:05.237334Z
A source-named dated measurement, never combined with another source.
Source: cited_works
62 of 62 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 69465d34-1e91-4d0f-8be4-fcbe807e7c05 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models LLaVA-OneVision-1.5: Fully Open Framework for Democratized Multimodal Training
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation edcd7512-d4b4-45b2-9d2a-0778e3952b64 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Qwen3-VL Technical Report
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c6f416eb-220a-4221-86e9-a94bf855832b · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc8738d5-5022-437b-ac97-d5cc843b92b1 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models LLaDA2.0: Scaling Up Diffusion Language Models to 100B
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 60a76bbc-69c0-44e6-ab0f-0fcb8e0ae680 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models SAM 3: Segment Anything with Concepts
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation f074caa9-2d56-40fd-9007-920e4ba0fb36 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Shikra: Unleashing Multimodal LLM's Referential Dialogue Magic
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 81a74e3f-d806-4c02-a06d-66b480a1592c · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Are we on the right way for evaluating large vision-language models? InNeurIPS, 2024
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32a425e3-c04f-4d7f-b7f3-01c3542cbf20 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 69c4fb90-6787-46bb-8740-8e444176e815 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Sdar: A syn- ergistic diffusion-autoregression paradigm for scalable sequence generation.arXiv preprint arXiv:2510.06303
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1034c832-6371-4d35-bb79-82a7104ad9f3 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Coconut: Modernizing coco segmentation
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1bcbb7fb-6833-411b-bd3e-a0fb4fd980d2 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models VLMevalKit: An open-source toolkit for evaluating large multi-modality models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7065d67-8233-4124-bfdf-2f89ce39d9a4 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Blink: Multimodal large language models can see but not perceive
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 985f5488-3e60-48a4-be6e-710af8944d4d · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models The Llama 3 Herd of Models
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6addfe4d-986b-410d-ae39-3528dbdf7ee1 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Dataseg: Taming a universal multi-dataset multi-task segmentation model
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4092cf1b-cc5e-4ae6-84ae-ec7ac589face · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8ea77c65-c133-45dc-a0e2-372d83264c7d · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation c9ff9b25-681c-4432-8761-78883a39ac0d · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Ai2d: A dataset for diagram understanding
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed7700ed-c843-4e75-8f84-02553ee9164b · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Segment anything
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7c1e478e-961e-4009-9c0d-72a8061d46e4 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models The scalability of simplicity: Empirical analysis of vision-language learning with a single transformer
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caf6e1e2-57b0-41ee-a72e-d0e331a3545a · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Seed-bench: Benchmarking multimodal large language models
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f702793-1820-4b27-9119-09987a479169 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Llava-onevision: Easy visual task transfer.TMLR, 2025
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6d5d9761-0f57-4a1e-915f-e3d09818a77e · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation eabc9dd0-9e18-4a5b-8158-9306d1033461 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Describe anything: Detailed localized image and video captioning
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d98e4add-d5db-43f3-ab25-04f496bbff13 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Microsoft coco: Common objects in context
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf609126-ac93-4c8d-a453-f751e9d57b4e · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Visual instruction tuning.NeurIPS, 2023
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57699af0-374f-489f-8875-5c7dc83de466 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Mmbench: Is your multi-modal model an all-around player? InEuropean conference on computer vision, pp
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 27bb208d-a5a3-46eb-b473-0b226e5eb6ff · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3eec8668-8fcb-4299-adf3-82246b1b453e · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Chartqa: A benchmark for question answering about charts with visual and logical reasoning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 75201859-afcf-43f4-83c6-887f0feccea7 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5a9c32be-0c02-4617-93f7-9729fd6f1cfb · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models InfographicVQA
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9fced0f7-1ff4-4aa8-a8ac-d5d7cf8e1356 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Open-o3 video: Grounded video reasoning with explicit spatio-temporal evidence
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation bb3416f3-4bd8-4940-9bf2-10cd1ddf524f · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models The flexibility trap: Why arbitrary order limits reasoning potential in diffusion language models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1a5f57a9-298c-47bc-9f48-08a33bf5db61 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Large language diffusion models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 365bd1b4-dc9e-4c97-9c5d-4b864b82fe2b · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Openai-gpt-5.2.https://openai.com/index/introducing-gpt-5-2/, 2025
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 961d46b0-bd76-4b90-bc32-731911991abf · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models d3LLM: Ultra-Fast Diffusion LLM using Pseudo-Trajectory Distillation
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation b69a7643-0ab1-49ae-95e6-8e99ba9cab11 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Glamm: Pixel grounding large multimodal model
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 40e00f14-7ed4-4f88-a8ec-68372a8c307b · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Objects365: A large-scale, high-quality dataset for object detection
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0262fa8a-2bb8-4cd6-85af-93a6e9077d2d · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Cambrian-1: A fully open, vision-centric exploration of multimodal llms
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7b3efd9e-afbd-4593-bb19-9b4516a6026e · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 599b1598-a3f2-4c7c-9796-fc209c233e0b · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models SigLIP 2: Multilingual Vision-Language Encoders with Improved Semantic Understanding, Localization, and Dense Features
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 376d24da-6733-471c-8065-713f086183e2 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Grasp any region: Towards precise, contextual pixel understanding for multimodal llms.ArXiv, abs/2510.18876
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5f12108a-f962-4534-ab2c-fe8eaa15d317 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Ross3d: Reconstructive visual instruction tuning with 3d-awareness
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b0bea613-3961-48fc-b414-37476d1e0217 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models V*: Guided visual search as a core mechanism in multimodal llms
Reference 43
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3c5f94fe-7cf3-4d5c-a796-4d8d0b1f3f78 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Realworldqa: A benchmark for real-world spatial understanding
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 631f3465-7d80-4135-9859-7c214aea7b70 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Lumina-dimoo: An omni diffusion large language model for multi-modal generation and understanding
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1f1d4d8b-b552-4b90-b73f-f5929da88acc · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Qwen3 Technical Report
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5526169c-264e-4da9-9212-d3e6b66e7435 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 9722ad0e-3013-4482-8184-58930a299833 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Mmada: Multimodal large diffusion language models
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 759bbd20-b2e3-4749-8f45-d7100e30db65 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Dream 7B: Diffusion Large Language Models
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d13cc417-d2c5-4d66-8cfd-cc2f299eca4d · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models LLaDA-V: Large Language Diffusion Models with Visual Instruction Tuning
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7e4f28f1-2df0-4fcd-a6fe-9c177e3daa78 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Sa2VA: Marrying SAM2 with LLaVA for Dense Grounded Understanding of Images and Videos
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2e7bb20e-13d2-4162-98d5-e1c8b9b0587f · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Visual reasoning tracer: Object-level grounded reasoning benchmark
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5ced1f86-011a-4c18-a68c-35faa9a26abe · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models arXiv preprint arXiv:2510.23603 , year=
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d8e3b033-faad-42a5-9651-3be9e6698292 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 54
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation aa64f7c4-f7a0-4cd7-ac01-3ea84e42919b · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models MMMU: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 55
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation db0bbb23-5a06-4e20-a995-698d56181b07 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? In ECCV, 2024
Reference 56
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8965d514-eac4-4041-a629-53856b3c9536 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models GPT4RoI: Instruction Tuning Large Language Model on Region-of-Interest
Reference 57
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 371d1551-fea6-4dd8-a715-6250d03aa42a · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Omg-llava: Bridging image-level, object-level, pixel-level reasoning and understanding.NeurIPS, 2024
Reference 58
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf4aeb30-1799-4ff3-a2a5-7e0bbab87631 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Pixel-SAIL: Single Transformer For Pixel-Grounded Understanding
Reference 59
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d4450f81-72ff-4bb2-8715-2f4d3a6ffcdc · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Bee: A high-quality corpus and full-stack suite to unlock advanced fully open mllms
Reference 60
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 6c463358-a249-4d29-8403-dc4d0dfafc75 · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models Samtok: Representing any mask with two words.arXiv preprint arXiv:2601.16093, 2026
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 7326d3a8-24dd-42ed-9d4a-1d90b0edbc3e · outbound
PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 62
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dc62360a-eee7-4d13-8555-e62edb66b550 · inbound
Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO PerceptionDLM: Parallel Region Perception with Multimodal Diffusion Language Models
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.