Pith. sign in

Paper Citation Record · LEDGER

Towards Multimodal In-Context Learning for Vision & Language Models

As of 8 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 5 inbound Pith citation observations for arXiv:2403.12736.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2403.12736 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 5 of 5 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:29:52.146055Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-10T13:45:28.217864Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 5d7fb8b5-7c7b-4f08-8ac1-8b13bd9c29eb · inbound

HAIBU-ReMUD: Reasoning Multimodal Ultrasound Dataset and Model Bridging to General Specific Domains cites this paper.

HAIBU-ReMUD: Reasoning Multimodal Ultrasound Dataset and Model Bridging to General Specific Domains Towards Multimodal In-Context Learning for Vision & Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:29:52.146055Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:29:52.146055Z digest=sha256:e1ebc60574226d1bf34ed6914271cd1080bc4b109a456c01638793d7315c06a1

Observation da1fab29-7ff0-4d7d-800c-e27fe95aa35b · inbound

True Multimodal In-Context Learning Needs Attention to the Visual Context cites this paper.

True Multimodal In-Context Learning Needs Attention to the Visual Context Towards Multimodal In-Context Learning for Vision & Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:33.800157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:33.800157Z digest=sha256:8068ba9dd69813ff4b9fbb7f21a0aad4add4d9ae6367604e4ba7f29bd442780d

Observation 7eeb1dd7-260b-4169-8e99-a7e472d65e32 · inbound

Generalizable Object Re-Identification via Visual In-Context Prompting cites this paper.

Generalizable Object Re-Identification via Visual In-Context Prompting Towards Multimodal In-Context Learning for Vision & Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-05T14:33:22.084330Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T14:33:22.084330Z digest=sha256:a0e827e9da86e716ccaca4b2a323558682208785fed08dbc825aad5bc2e2cb10

Observation dcf3cc4d-844a-41ac-a697-cb65eb5830f6 · inbound

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks cites this paper.

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks Towards Multimodal In-Context Learning for Vision & Language Models

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:45:28.220406Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=arxiv_source observed=2026-05-10T13:41:37.942145Z digest=sha256:e64e3cbdce58105ccae299682bef69b0e3ae4c215981452f967406003f93da58

Observation ea5853cc-6f27-4965-8e07-89a284b14132 · inbound

BanglaWild: An In-the-Wild Bengali Scene Text Recognition Benchmark for OCR and Vision-Language Models cites this paper.

BanglaWild: An In-the-Wild Bengali Scene Text Recognition Benchmark for OCR and Vision-Language Models Towards Multimodal In-Context Learning for Vision & Language Models

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-05T10:30:24.277663Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T10:30:24.277663Z digest=sha256:898c5a8423bac93c1e16862e3fff488e955e2ddd5aa880824aff573624851ea8