Pith. sign in

Paper Citation Record · LEDGER

Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2403.20271.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.20271 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T13:45:55.390063Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.285599Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5c953bd2-e379-4792-a718-1b9d1fd8a12d · inbound

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding cites this paper.

VideoLLaMA 3: Frontier Multimodal Foundation Models for Image and Video Understanding Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 78

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:19:59.969643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:19:59.603343Z digest=sha256:9ecc15914db77a61a67e41a7f73f5728af76b18deb0e9fd6f22859cd5ce1f81b

Observation 1ebdb2ac-b14b-4722-991f-169e278f991f · inbound

LPOI: Listwise Preference Optimization for Vision Language Models cites this paper.

LPOI: Listwise Preference Optimization for Vision Language Models Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T13:45:55.390063Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T13:45:55.390063Z digest=sha256:3e826f3c64c6596cdda806995599537caa14750e3a7fed2d529d7dc105f3af43

Observation 8038ddba-2872-40b6-9f42-843263afc2fe · inbound

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs cites this paper.

Zooming from Context to Cue: Hierarchical Preference Optimization for Multi-Image MLLMs Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-07T13:14:09.279344Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:14:09.279344Z digest=sha256:d790b99b9cd050484be8708786f6316c02149ba64c1087a731a7a6bc911fe26b

Observation 8aa6eeb3-2245-4d99-8f1c-4cef689c819a · inbound

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought cites this paper.

RSVP: Reasoning Segmentation via Visual Prompting and Multi-modal Chain-of-Thought Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T11:08:42.177990Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T11:08:42.177990Z digest=sha256:806b7836171c4a91ebeb02ec201b2684d1d4103ebf72f297f480c7a5d6c70f4b

Observation 734693f4-cdbe-494d-a851-2c3d8c130a24 · inbound

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos cites this paper.

Perceive Anything: Recognize, Explain, Caption, and Segment Anything in Images and Videos Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T10:28:11.715662Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:28:11.715662Z digest=sha256:9c583bf2b2785d6298d2715a2b3870ff54eedc6a47fffb07be1cb3a796e207af

Observation 88536864-af88-40b0-8df6-7f58519084cf · inbound

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World cites this paper.

DenseWorld-1M: Towards Detailed Dense Grounded Caption in the Real World Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-06T21:27:40.583636Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:27:40.583636Z digest=sha256:d27855dc2dd5e519821edc60941033a7ed26cb79044f8c3942d930f5da161263

Observation 13065974-bb15-4c14-9943-db6377dce38b · inbound

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details? cites this paper.

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details? Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-05T23:35:21.467830Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:35:21.467830Z digest=sha256:c5246f72b9b302f07c77e44545f7dce5b838b77dde016a2f31806a1b8f99210d

Observation 3f185046-efe2-4236-9ccb-86527044b2be · inbound

The high-speed X-ray camera on AXIS: design and performance updates cites this paper.

The high-speed X-ray camera on AXIS: design and performance updates Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-05T18:48:03.409499Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T18:48:03.409499Z digest=sha256:cbcc3428b6f4cfe6a6282b4f7681b28cac5101dcde35ce2e2ff2ab8ee371ef9b

Observation b819030b-abb2-474b-a43c-1fdb5303f7b7 · inbound

VoCap: Video Object Captioning and Segmentation from Any Prompt cites this paper.

VoCap: Video Object Captioning and Segmentation from Any Prompt Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-05T14:01:17.690922Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:01:17.690922Z digest=sha256:f606d01f1b79e4c53fe93914803e6c13a60c996828ac49fb4412fb60109594e1

Observation 3eb49cf9-3455-43e8-8fe9-ea4402bdc796 · inbound

Eevee: Towards Close-up High-resolution Video-based Virtual Try-on cites this paper.

Eevee: Towards Close-up High-resolution Video-based Virtual Try-on Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-17T06:39:10.444767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T06:34:45.045267Z digest=sha256:e0fecaa4f1f8f9851537af0c8dfab11d56e9ae1c0eaa1764ff5eeb3bad854c76

Observation a8333f38-25bb-48dd-b407-e8f5f10b7d6e · inbound

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition cites this paper.

MICo-150K: A Comprehensive Dataset Advancing Multi-Image Composition Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:21:23.357910Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T00:20:58.483350Z digest=sha256:8f427acf246c99931b4cd5f49fdd09fb7146285c5370f524c0c189b46ff4f2dd

Observation 84cc4755-f6aa-404c-a487-402e84ec102e · inbound

VABench: A Comprehensive Benchmark for Audio-Video Generation cites this paper.

VABench: A Comprehensive Benchmark for Audio-Video Generation Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-17T00:08:43.880478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-17T00:03:45.576961Z digest=sha256:3233c415665b7398702cd5024d0624f749b67272f40fad24bf97418acef70e57

Observation 099aca5f-e86d-4736-b73d-a26ef6ccb138 · inbound

Enhancing Foundation VLM Robustness to Missing Modality: Scalable Diffusion for Bi-directional Feature Restoration cites this paper.

Enhancing Foundation VLM Robustness to Missing Modality: Scalable Diffusion for Bi-directional Feature Restoration Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 34

Resolution
verified exact
arxiv_id, observed 2026-05-16T08:37:37.247123Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-16T08:33:49.841678Z digest=sha256:bc50ab6cd36ee546ccf45e7cd2363d0f35e4fa0c5aee9f347b2d62f7107d1505

Observation bc407873-50c7-4d54-a77b-68db0bdaf033 · inbound

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes cites this paper.

E-VLA: Event-Augmented Vision-Language-Action Model for Dark and Blurred Scenes Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-13T09:40:35.631188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T09:40:35.631188Z digest=sha256:bc406b23bdec7ee855789cc9cc49a03b1a8270768b5cce330c207d9b5eac7b15

Observation 08b66f69-779f-4a4c-84da-280aaae95099 · inbound

Less Detail, Better Answers: Degradation-Driven Prompting for VQA cites this paper.

Less Detail, Better Answers: Degradation-Driven Prompting for VQA Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 3

Resolution
verified exact
arxiv_id, observed 2026-05-10T22:05:47.909302Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T20:17:01.867903Z digest=sha256:c2ebae55c4d4b8a72c06034d7d2caba74ecf1029a86bc83c7bad38acdbb07c08

Observation 831e340b-8378-43ec-b7c0-8f747c970060 · inbound

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation cites this paper.

LMMs Meet Object-Centric Vision: Understanding, Segmentation, Editing and Generation Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 95

Resolution
verified exact
arxiv_id, observed 2026-05-11T10:11:11.116785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T15:35:37.095627Z digest=sha256:67ce872c1d6d92ee6dbea192698e8e6aaeef95774378bda932cf13e6d89fd701

Observation b2302e69-4835-4f13-a134-b540e1896a1c · inbound

WOW-Seg: A Word-free Open World Segmentation Model cites this paper.

WOW-Seg: A Word-free Open World Segmentation Model Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-19T21:27:48.040652Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T21:23:14.311122Z digest=sha256:ee3fd946d547547451d793fdec192f5c982988324a8f61ed27b6d35887827578

Observation ea1e2af6-8aae-4be1-b89c-5dd4ddf80e68 · inbound

LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding cites this paper.

LocateAnything: Fast and High-Quality Vision-Language Grounding with Parallel Box Decoding Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 2

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T17:53:46.750455Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-29T17:52:59.346637Z digest=sha256:b43aa8615cc041e7b38c436072b2f75b4058de7bc0a18f13d7d91e248e211623

Observation 412f60c4-fc1f-48c4-9595-47af92b3b169 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 35

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.287671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:8dc6ebc575d5972dcfe0d4f549d395427985c50c79b5bf5f121955f1acf7ce64

Observation 9ceaa00c-f8bb-40b8-93d5-11c989a28ae9 · inbound

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO cites this paper.

Actor as Its Own Critic: Unifying Region Understanding and Localization via CycleGRPO Draw-and-Understand: Leveraging Visual Prompts to Enable MLLMs to Comprehend What You Want

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-14T04:38:05.237334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T04:38:05.237334Z digest=sha256:3380e652460813324ba4ee24d0f9f9220d2cc4b766642e7680a9638f490a51a4