Pith. sign in

Paper Citation Record · LEDGER

Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 6 inbound Pith citation observations for arXiv:2410.21220.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2410.21220 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 6 of 6 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 6 of 6 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T04:57:54.119975Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-16T15:27:04.356488Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 366a2314-cb85-48de-a10c-ec5e863ee7ff · inbound

Agent-based Condition Monitoring Assistance with Multimodal Industrial Database Retrieval Augmented Generation cites this paper.

Agent-based Condition Monitoring Assistance with Multimodal Industrial Database Retrieval Augmented Generation Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-07T04:57:54.119975Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T04:57:54.119975Z digest=sha256:18cb90db460f24f6eb3e69344a429faaaeaded956f2c02000cc827f2f6486daf

Observation fdd8921a-fe0a-40e5-b704-e73f448640eb · inbound

MMSearch-R1: Incentivizing LMMs to Search cites this paper.

MMSearch-R1: Incentivizing LMMs to Search Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-16T15:27:04.359334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-16T15:27:04.228144Z digest=sha256:8b17670980ae5fa94a5fc9e39970b5b44247dfcaae4fe349b0292edf871d6fb8

Observation fdafa1c5-c530-4513-85cb-d8b059f3caab · inbound

Hidden Tail: Adversarial Image Causing Stealthy Resource Consumption in Vision-Language Models cites this paper.

Hidden Tail: Adversarial Image Causing Stealthy Resource Consumption in Vision-Language Models Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-05T16:15:28.012224Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T16:15:28.012224Z digest=sha256:3163cdd6b4df96482d8a35c8450df6890803179fe85ed9ec1a0c98035f298331

Observation 87c3752c-9c26-4dbd-a70f-c555aeab7736 · inbound

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents cites this paper.

DR-MMSearchAgent: Deepening Reasoning in Multimodal Search Agents Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines

Reference 65

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T12:31:07.876274Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T03:21:30.732925Z digest=sha256:540dcf8671e443958e6a46591c51fafd745664196ddd74f256caf7b1827a9cbe

Observation eb930070-8b28-49f4-8e0b-cc3f4616b474 · inbound

ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards cites this paper.

ProMMSearchAgent: A Generalizable Multimodal Search Agent Trained with Process-Oriented Rewards Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:41:10.100388Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-10T01:12:17.469552Z digest=sha256:6672cb2d460a87e6ff5db2afcd266a3fd793748bd8e9c6dd46d48109b591cd5e

Observation 69f8f00b-ff32-477a-a004-e9721ce81336 · inbound

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG cites this paper.

Reason Before You Retrieve: Agentic Planning for Multi-modal RAG Vision Search Assistant: Empower Vision-Language Models as Multimodal Search Engines

Reference 156

Resolution
unresolved
no resolver link, observed 2026-08-02T10:20:56.272120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-02T10:20:56.272120Z digest=sha256:a8cb61fa545338ecc76f8bd3e2cdb6e1ac2ef8e7831720d2ba60b89b4107e0a0