Pith. sign in

Paper Citation Record · LEDGER

Many-Shot In-Context Learning in Multimodal Foundation Models

As of 22 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 20 inbound Pith citation observations for arXiv:2405.09798.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2405.09798 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 20 of 20 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:05:19.268015Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-25T04:25:19.781518Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation 8689c888-24a6-4676-8f49-0fd6649775e0 · inbound

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding cites this paper.

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T01:09:30.475892Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-17T01:09:30.360275Z digest=sha256:a4ca528a415075c94a2081fbac076fff76276df113072d937fb4a141bdd9b3ad

Observation c324eeda-50c0-4d0a-a6d4-63947787e77b · inbound

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models cites this paper.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T06:01:53.841270Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:ab262439e80c5af57b606afa6b7e8509b12a987e9436b32d1fd5ad1fc7c1502b

Observation bda496a3-7261-4297-a81d-5ad25e3707ea · inbound

LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations cites this paper.

LMAct: A Benchmark for In-Context Imitation Learning with Long Multimodal Demonstrations Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-12T04:25:49.796437Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-12T04:25:49.796437Z digest=sha256:cabbc608a95907c687ef43b49f978399e85c1d4c813534549b4124bd8677eb6c

Observation 7934e86c-77b5-41b4-9c21-78f8b1f5b949 · inbound

Align, Generate, Learn: A Novel Closed-Loop Framework for Cross-Lingual In-Context Learning cites this paper.

Align, Generate, Learn: A Novel Closed-Loop Framework for Cross-Lingual In-Context Learning Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-11T17:25:18.865670Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T17:25:18.865670Z digest=sha256:379539696d57e187961d39883f0b2da6b27ea120ce18d892fa684f98c3c6895d

Observation 5ab54576-f894-4307-a2c7-413cd260af69 · inbound

Error-driven Data-efficient Large Multimodal Model Tuning cites this paper.

Error-driven Data-efficient Large Multimodal Model Tuning Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-11T11:18:31.447699Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-11T11:18:31.447699Z digest=sha256:91600430b1a7aa5222c089b19603e1485d227c88602e47747fdfbb473ecf0eb1

Observation de007016-4963-4ae8-94f5-8b7ecf2e55b7 · inbound

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey cites this paper.

Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 185

Resolution
unresolved
no resolver link, observed 2026-08-11T14:59:01.959621Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-11T14:59:01.959621Z digest=sha256:bcb295fd45b53e8ed9d267bcbf4faf44145aabffb5363d4f9e7809ab2a95a74e

Observation 9b2465b4-c836-4a5e-9a6b-bb7a62a0222d · inbound

Visual RAG: Expanding MLLM visual knowledge without fine-tuning cites this paper.

Visual RAG: Expanding MLLM visual knowledge without fine-tuning Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-10T18:59:22.362716Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T18:59:22.362716Z digest=sha256:b140594f15da33ac2958437ee779e722935f0ed7b24d535d24f794f828c76e33

Observation 566c462b-768c-4f74-92e5-1c45327b56d5 · inbound

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents cites this paper.

MedAgentBench: A Realistic Virtual EHR Environment to Benchmark Medical LLM Agents Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-10T14:58:40.203221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T14:58:40.203221Z digest=sha256:11bf7d7b0fa042245c0d88307f4ae1636d399c9a0aee3a4cdbaa8ce973ea7f47

Observation 83c0b738-d009-48d4-87bd-3ad0b198b3f7 · inbound

From Few to Many: Self-Improving Many-Shot Reasoners Through Iterative Optimization and Generation cites this paper.

From Few to Many: Self-Improving Many-Shot Reasoners Through Iterative Optimization and Generation Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-09T19:31:49.650798Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-09T19:31:49.650798Z digest=sha256:08198ea57178066f197b8ffe39b8b1eb5e489c823fb80b5c7f7355867cff3bf2

Observation 41d89f27-34a1-4bbc-9c4f-63c89b39b4f0 · inbound

Selecting Demonstrations for Many-Shot In-Context Learning via Gradient Matching cites this paper.

Selecting Demonstrations for Many-Shot In-Context Learning via Gradient Matching Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T10:45:25.419580Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T10:45:25.419580Z digest=sha256:fa34ea0668a246881cbd537e74cca8f70e4fd380ec277db5d32080633805a103

Observation 973035ca-77dd-43cd-8391-1e740949f05d · inbound

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks cites this paper.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 31

Resolution
metadata mismatch
arxiv_id, observed 2026-05-19T05:57:08.083275Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:8e49280947597df519763c6fa38233703c90a71e261801b852893f8ba64f8b7e

Observation 81541ab2-1b20-4b79-8230-ef2d6bf3d74c · inbound

True Multimodal In-Context Learning Needs Attention to the Visual Context cites this paper.

True Multimodal In-Context Learning Needs Attention to the Visual Context Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.320527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.320527Z digest=sha256:f62917b6502cfd98148b5b5a7e729371508053e517be97ce2973c47050fc39bc

Observation f949f0e5-46a8-445c-be4f-4a7c575cc8a6 · inbound

True Multimodal In-Context Learning Needs Attention to the Visual Context cites this paper.

True Multimodal In-Context Learning Needs Attention to the Visual Context Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.377976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.377976Z digest=sha256:52e76727a2fcb7d18bc8fd7fb02b4db17bc4dd9fab2eeefb6878d2d399ddb6af

Observation 95eb9f2f-dd60-4303-995a-fb97ec87b6d8 · inbound

Towards Compute-Optimal Many-Shot In-Context Learning cites this paper.

Towards Compute-Optimal Many-Shot In-Context Learning Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:20:38.661092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:20:38.661092Z digest=sha256:132fd86b47077490d7722fb1aa95b5f367f1e1bbe123cb100a92894de13d66e7

Observation a01c224e-b043-443d-9d2f-17963259643c · inbound

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors cites this paper.

Visual Species Recognition with Large Multimodal Models as Post-Hoc Correctors Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T17:20:52.476377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T17:20:52.476377Z digest=sha256:d877c3742d03f8993bec8b0958fcb6860c566df53d2d9827047ea08416cf74b3

Observation c6c0965f-1f6e-458e-9f0e-fd5d4f94c0b7 · inbound

Personal Visual Context Learning in Large Multimodal Models cites this paper.

Personal Visual Context Learning in Large Multimodal Models Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-12T07:06:37.577186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-12T03:42:15.402131Z digest=sha256:f94ca16006823143aa67b693c71e0bf3b8e7b92824083f8a164ab5e8f893429a

Observation c5992951-048a-4136-8f99-84469c93f655 · inbound

When Youth Enter the Algorithmic Wild: Discovering and Understanding Potentially Harmful Teen Videos on Douyin and Kwai cites this paper.

When Youth Enter the Algorithmic Wild: Discovering and Understanding Potentially Harmful Teen Videos on Douyin and Kwai Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 22

Resolution
verified exact
arxiv_id, observed 2026-05-25T04:25:19.783887Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-05-25T04:20:27.456558Z digest=sha256:2a7ab45b17d17496bb388b19b3c997208533f3baa7e5363ed26dcf8e4e6dc98c

Observation 610dae7e-06c2-4739-9085-c936289a3846 · inbound

In-Context Learning for Wound Classification with Small Multimodal Language Models cites this paper.

In-Context Learning for Wound Classification with Small Multimodal Language Models Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-01T14:15:59.840877Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T14:15:59.840877Z digest=sha256:b3c02b3fa2e0f4630ce58cf4b59a8aa6f3e1b2411a1d44d4fd92c3a00cfea76b

Observation 081ba260-c89f-4407-88bd-80a52c4d203e · inbound

Inference-Time Agentic Decision Rules Beat Longer Evolving Search for Multi-Image Medical Reasoning cites this paper.

Inference-Time Agentic Decision Rules Beat Longer Evolving Search for Multi-Image Medical Reasoning Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-01T05:38:56.492179Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T05:38:56.492179Z digest=sha256:f6eb1e7fc04712e78b35db0b5f4d075477a1f8a5b1e2c47255d734b1de8a24bd

Observation 6b08c0e2-cb6d-4a25-a925-a9af27946d64 · inbound

In-Context Collapse in Vision-Language Models and How to Mitigate it? cites this paper.

In-Context Collapse in Vision-Language Models and How to Mitigate it? Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T15:05:19.268015Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:05:19.268015Z digest=sha256:263f1840327ac6b4e6304f95f99cf455751538f3f20e547ec2055f6bd991c26e