Pith. sign in

Paper Citation Record · LEDGER

Visual Prompting in Multimodal Large Language Models: A Survey

As of 7 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 16 inbound Pith citation observations for arXiv:2409.15310.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2409.15310 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 16 of 16 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 16 of 16 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:36:15.716207Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T16:03:08.102407Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 7aaf53df-d37c-4d31-b5c1-7996665ae2f2 · inbound

Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models cites this paper.

Uncertainty-o: One Model-agnostic Framework for Unveiling Uncertainty in Large Multimodal Models Visual Prompting in Multimodal Large Language Models: A Survey

Reference 66

Resolution
unresolved
no resolver link, observed 2026-08-07T05:36:15.716207Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:36:15.716207Z digest=sha256:717376ac710640a57cd1fcf53b7a0509fcc627d72c349d0ca34ce1e3b884567c

Observation ee4b47a9-2a6b-4935-b17b-3a2656142bf0 · inbound

AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs cites this paper.

AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs Visual Prompting in Multimodal Large Language Models: A Survey

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:47.733190Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:47.733190Z digest=sha256:0c731c9a5cb121a380c2535332de180d35d760016f4684cb23436c13923b8d5a

Observation 8f1ad2cc-cb70-4d95-a6eb-141ee985027d · inbound

PDB-Eval: An Evaluation of Large Multimodal Models for Description and Explanation of Personalized Driving Behavior cites this paper.

PDB-Eval: An Evaluation of Large Multimodal Models for Description and Explanation of Personalized Driving Behavior Visual Prompting in Multimodal Large Language Models: A Survey

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T14:36:55.827401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T14:36:55.827401Z digest=sha256:0f69230dce48ed80d93391339265e47c6796e9a2221166c48b645868a89bd3dc

Observation d7e906a8-9839-4d82-9328-c67c32ea5419 · inbound

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details? cites this paper.

Finding Needles in Images: Can Multimodal LLMs Locate Fine Details? Visual Prompting in Multimodal Large Language Models: A Survey

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-05T23:35:24.104443Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:35:24.104443Z digest=sha256:72b596b8f9968ef0c2e038ddaf6a8efd97abdf16713c52a8363572ee1faf644e

Observation 3e86d794-ff2a-413c-b997-c1cb4c560696 · inbound

Invisible Watermarks, Visible Gains: Steering Machine Unlearning with Bi-Level Watermarking Design cites this paper.

Invisible Watermarks, Visible Gains: Steering Machine Unlearning with Bi-Level Watermarking Design Visual Prompting in Multimodal Large Language Models: A Survey

Reference 85

Resolution
unresolved
no resolver link, observed 2026-08-05T21:01:26.976295Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T21:01:26.976295Z digest=sha256:799a2d8edc1c4a60636aa977959016a2fc3a4eb3fc200f66b5563673a89afc8e

Observation 45f422a6-bf94-4aba-a52a-2f60719e77aa · inbound

A Survey on Video Temporal Grounding with Multimodal Large Language Model cites this paper.

A Survey on Video Temporal Grounding with Multimodal Large Language Model Visual Prompting in Multimodal Large Language Models: A Survey

Reference 87

Resolution
unresolved
no resolver link, observed 2026-08-05T23:32:17.850843Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T23:32:17.850843Z digest=sha256:762d7d3365004acbe23a19ef9118c2b088209f95c1991179fb9e6993d6530d96

Observation be421962-57d0-4d6b-a57b-c015f11c8160 · inbound

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs cites this paper.

Dictionary-Aligned Concept Control for Safeguarding Multimodal LLMs Visual Prompting in Multimodal Large Language Models: A Survey

Reference 113

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:35:57.458742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T18:04:05.157103Z digest=sha256:bf29ed1a599d9441381a42be44f16411ca18ed90d5a6e275b1bcff2112b246d0

Observation 6005aaec-6a6e-487f-9189-3de52d2a65b5 · inbound

STORM: End-to-End Referring Multi-Object Tracking in Videos cites this paper.

STORM: End-to-End Referring Multi-Object Tracking in Videos Visual Prompting in Multimodal Large Language Models: A Survey

Reference 74

Resolution
verified exact
arxiv_id, observed 2026-05-11T08:56:00.442212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T16:25:31.777907Z digest=sha256:21659757b045daa4b17b7c44608f68a6129e7c45c1186c8e6e4d618bdf37c09e

Observation 32c52558-2f0c-44c1-9586-f9c50a23884d · inbound

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification cites this paper.

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification Visual Prompting in Multimodal Large Language Models: A Survey

Reference 50

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T06:56:47.420763Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-10T06:55:33.464951Z digest=sha256:aef54bdf617430047a3cb3a1823cdd754d281a09f3be6e5d425721e04add7434

Observation 3f808a75-1b6c-497b-ab1a-dbc0d9006b24 · inbound

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification cites this paper.

Rule-VLN: Bridging Perception and Compliance via Semantic Reasoning and Geometric Rectification Visual Prompting in Multimodal Large Language Models: A Survey

Reference 52

Resolution
unresolved
no resolver link, observed 2026-07-12T19:10:41.883296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T19:10:41.883296Z digest=sha256:129aa69c4fefe99699c694c1625632000da0003ea1a68f645bf6b3b71fcadf78

Observation c2e6a976-bb1d-443f-89b0-e8f8e9a0a29f · inbound

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning cites this paper.

Generate, Filter, Control, Replay: A Comprehensive Survey of Rollout Strategies for LLM Reinforcement Learning Visual Prompting in Multimodal Large Language Models: A Survey

Reference 137

Resolution
verified exact
arxiv_id, observed 2026-05-10T23:15:49.439868Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T19:15:27.406778Z digest=sha256:9c4023c72626108a13d2fc911f6b6e9a008b603741948fbc2bb537b0c929b178

Observation 5a3cfb1b-a25e-4961-b5b8-646f10e721cf · inbound

Skill-CMIB: Multimodal Agent Skill for Consistent Action via Conditional Multimodal Information Bottleneck cites this paper.

Skill-CMIB: Multimodal Agent Skill for Consistent Action via Conditional Multimodal Information Bottleneck Visual Prompting in Multimodal Large Language Models: A Survey

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T08:01:28.145905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-12T01:24:59.721212Z digest=sha256:ddc6f4423d914a4213360f92a8451ce3e38765093a1dce41ec7f380bb4cdc013

Observation 1a7e3a96-1854-41d2-bc00-f77617cb93cc · inbound

OLIVIA: Online Learning via Inference-time Action Adaptation for Decision Making in LLM ReAct Agents cites this paper.

OLIVIA: Online Learning via Inference-time Action Adaptation for Decision Making in LLM ReAct Agents Visual Prompting in Multimodal Large Language Models: A Survey

Reference 86

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T02:27:07.225834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-13T02:24:42.530103Z digest=sha256:c080da079c18c26af59e6dbf54798f9b55a51b1cc66f4893a2628d17f9af3a4f

Observation 9d51bcf7-f6d3-43cc-9d9d-296b3341e32c · inbound

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding cites this paper.

Minerva-Ego: Spatiotemporal Hints for Egocentric Video Understanding Visual Prompting in Multimodal Large Language Models: A Survey

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-19T16:03:08.104159Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-19T16:02:53.887605Z digest=sha256:138a3cb65b0ffae146c25670480059993228a21c4c6215365cef870d2d29d2d6

Observation e0630495-5989-46a3-9cb4-bb265b6a0e85 · inbound

Plover: Steering GUI Agents through Plan-Centric Interaction cites this paper.

Plover: Steering GUI Agents through Plan-Centric Interaction Visual Prompting in Multimodal Large Language Models: A Survey

Reference 75

Resolution
unresolved
no resolver link, observed 2026-08-01T23:57:13.850939Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T23:57:13.850939Z digest=sha256:3d6d83179902c352a06191a6538e1b1d9d970d6680849b3630e423226791ba33

Observation c9cbef5a-a706-4849-baec-6e32fb03b999 · inbound

Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs cites this paper.

Seeing What Is Actually There: PriVE-Bench and PriVE-Tools for Counterfactual Evaluation of Agentic Visual Evidence in VLMs Visual Prompting in Multimodal Large Language Models: A Survey

Reference 54

Resolution
unresolved
no resolver link, observed 2026-08-02T06:18:16.279690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T06:18:16.279690Z digest=sha256:a80dd1385e50bc8d1ea3ce3f0ba62a972880cc36eb971075c087413dace9d11d