Pith. sign in

Paper Citation Record · LEDGER

Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2404.04514.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.04514 v1

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T12:02:10.543405Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T15:09:55.132003Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation d4f1458c-9434-43a7-96cd-a74cffb507a2 · inbound

HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models cites this paper.

HSCR: Hierarchical Self-Contrastive Rewarding for Aligning Medical Vision Language Models Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T12:00:48.850779Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:00:48.850779Z digest=sha256:a29ee5c551e50e95856e1e59a95a9adb2b9b11fd3746b128d2530df5796978a5

Observation a06393d7-30f2-4f3a-af02-9e35869be43a · inbound

Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering cites this paper.

Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T12:02:10.543405Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:02:10.543405Z digest=sha256:3b7bb244615223ec33e5ed84d12a99d5a72e066668ecebeb006aad488c2e18fa

Observation b698e741-5e69-45fe-b013-9b060e96ccaa · inbound

AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs cites this paper.

AutoV: Loss-Oriented Ranking for Visual Prompt Retrieval in LVLMs Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T23:49:44.479493Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:49:44.479493Z digest=sha256:b8c3e5d2ce9628135d644fa8e71d51d2b5e0e84bf55ea3be6a395795a2af4826

Observation a1ad24e4-5e56-4f70-b626-6dd19f423b2c · inbound

V2T-CoT: From Vision to Text Chain-of-Thought for Medical Reasoning and Diagnosis cites this paper.

V2T-CoT: From Vision to Text Chain-of-Thought for Medical Reasoning and Diagnosis Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T23:11:46.047700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T23:11:46.047700Z digest=sha256:a7bf833faecff2cc792c77604a957d15c4e64d0c51e65c50c7813c4aa9512b4d

Observation d373e16c-9b10-4279-bce9-c2ac6828bcd0 · inbound

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models cites this paper.

From Structure to Synergy: A Survey of Vision-Language Perception Paradigm Evolution in Multimodal Large Language Models Joint Visual and Text Prompting for Improved Object-Centric Perception with Multimodal Large Language Models

Reference 125

Resolution
verified exact
arxiv_id, observed 2026-07-04T15:09:55.133807Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-06-26T01:50:54.242508Z digest=sha256:b7b05c7a2978b4c8282092a221e22f66fd9b3311881aee45d0900318181d9545