Pith. sign in

Paper Citation Record · LEDGER

MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2309.07915.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2309.07915 v3

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00

measured 25 of 25 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-10T22:56:54.862312Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-03T11:28:04.249589Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation c003e649-d936-4595-ac39-8a0569c65b9f · inbound

Otter: A Multi-Modal Model with In-Context Instruction Tuning cites this paper.

Otter: A Multi-Modal Model with In-Context Instruction Tuning MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 104

Resolution
verified exact
arxiv_id, observed 2026-05-15T02:43:47.921205Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T02:43:47.775691Z digest=sha256:5131a7aa8107a15c191d98f235f8144720c9d141e870e7e45a4049d7d48d0219

Observation 49c71f5b-2ca1-4f26-ba5c-ee4fda747cc5 · inbound

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models cites this paper.

MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 61

Resolution
verified exact
arxiv_id, observed 2026-05-10T20:25:34.384358Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-10T20:25:33.854923Z digest=sha256:9ee591b100ebea7a76355182ad83278bbb2bc81885a6345ae302600fe967e27a

Observation c32ec879-219d-46c8-a17e-5888f035abde · inbound

A Survey on Multimodal Large Language Models cites this paper.

A Survey on Multimodal Large Language Models MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 178

Resolution
verified exact
arxiv_id, observed 2026-05-16T02:56:42.114548Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T02:56:41.658658Z digest=sha256:ab6b43409b8f930ba742eb6e3f2ecb4bc55c353565def5cc4d9755c20f145e20

Observation 8f8cd8e6-452f-4c3f-a7ac-9da087753530 · inbound

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition cites this paper.

InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 102

Resolution
verified exact
arxiv_id, observed 2026-05-17T13:48:48.730141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T13:48:48.661566Z digest=sha256:5de3c58deb797790f4e5f55576c07da0d1db9d8f1f5c216e578734bbcaebedc0

Observation 7b199355-444a-4547-a8df-919470dc953e · inbound

SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models cites this paper.

SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 42

Resolution
verified exact
arxiv_id, observed 2026-05-17T03:03:26.811111Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-17T03:03:26.723464Z digest=sha256:3077ebcf3b55257e08468c6b45374e2867188df4adc84903908a353acde6cc0a

Observation dff90377-1243-4301-8102-07998a1a663e · inbound

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI cites this paper.

MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 93

Resolution
verified exact
arxiv_id, observed 2026-05-15T05:37:41.773286Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-15T05:37:41.401736Z digest=sha256:5582befb6a2dccc143d126f0fd50f47bd775373f5c3b2e1da486ed89eec62f3f

Observation 82d29ec2-bf0b-4894-a809-45bbdbbc10e9 · inbound

AppAgent: Multimodal Agents as Smartphone Users cites this paper.

AppAgent: Multimodal Agents as Smartphone Users MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 46

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T10:16:43.863403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-17T10:16:43.364787Z digest=sha256:5d7f71e8fd87e135b604ee90b61f8e22720c26f2bd2c8f022f5a63bb081405f6

Observation 63612e85-d06c-44c0-85a0-51c889966839 · inbound

Hallucination of Multimodal Large Language Models: A Survey cites this paper.

Hallucination of Multimodal Large Language Models: A Survey MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 215

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:33:33.408364Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-11T12:33:32.631346Z digest=sha256:203a306cf76e8efebf82e0f7e1dc7142b248bfa0b8703eb9621929319a5a6dfb

Observation 873119cc-23fb-4666-8c34-0fbe0df37024 · inbound

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models cites this paper.

mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 270

Resolution
verified exact
arxiv_id, observed 2026-05-20T06:20:36.589141Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-20T06:20:36.235304Z digest=sha256:332e0b3bd9ea7bb304b406d9c2aee24d03ca591bd98913e89f5a66c73b4e447c

Observation 6e32069a-5927-46d7-87ed-ff66f9af7781 · inbound

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark cites this paper.

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-14T00:51:48.489537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-14T00:51:48.163349Z digest=sha256:b29497ee96b8caf8f6d67e612c07773592e07c9e665c2759c74ac95d4f0a0b24

Observation 5d7c42c9-d3fb-4109-9bfc-10a9d1ece2e5 · inbound

Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding cites this paper.

Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 62

Resolution
unresolved
no resolver link, observed 2026-08-10T22:56:54.862312Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-10T22:56:54.862312Z digest=sha256:c7529e17f79afd46e97ffe62fa05659a7ae114ef9ca11394d0715229366733d7

Observation 980f9d5a-59ed-459f-943b-52df29e9c27d · inbound

Efficiently Enhancing General Agents With Hierarchical-categorical Memory cites this paper.

Efficiently Enhancing General Agents With Hierarchical-categorical Memory MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T13:21:27.449934Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T13:21:27.449934Z digest=sha256:3c83ac99fd829b66b05702ccf4445004c5c87386ee3b712354000a19eb4bc932

Observation b7cadc4d-47b2-4570-921e-e97df1d5db9c · inbound

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs cites this paper.

LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 65

Resolution
unresolved
no resolver link, observed 2026-08-06T21:19:45.166758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:19:45.166758Z digest=sha256:cc883370066bc22dbaab1abb4c60ee3d319fa372ee6bcc5ef8e14ef530482df0

Observation 17b80f1f-639d-4ed0-8004-9db7e39d7fca · inbound

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs cites this paper.

DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 84

Resolution
unresolved
no resolver link, observed 2026-08-06T17:38:19.968980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T17:38:19.968980Z digest=sha256:7cd80b71680d652c265fe186bacca36ed0153a522c896379f5a0220e6b951a18

Observation 26378829-0ef0-4a8b-8169-bac28bf660ad · inbound

Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual Reasoning cites this paper.

Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual Reasoning MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 61

Resolution
unresolved
no resolver link, observed 2026-08-06T17:09:49.143377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T17:09:49.143377Z digest=sha256:ca10411ed412275db31909ec71ac8c3d017d9bf78e5d060d50dd88d9f1179094

Observation f304a5c6-2cb3-4968-9053-1d34374fe6bb · inbound

True Multimodal In-Context Learning Needs Attention to the Visual Context cites this paper.

True Multimodal In-Context Learning Needs Attention to the Visual Context MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:35.387308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:35.387308Z digest=sha256:66102e11f7b88b718bd067216500ae509a53e45aa9844b140469cc1353d57943

Observation dc0acddf-0452-4542-8916-34dbb5c29744 · inbound

Region-Level Context-Aware Multimodal Understanding cites this paper.

Region-Level Context-Aware Multimodal Understanding MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-05T19:39:48.064324Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T19:39:48.064324Z digest=sha256:26ecd1b99a2d6c5f800e896723b7ccdcf6915fd7f041ea0cfa33a73e723232db

Observation c849f00e-97e6-47c5-bef1-9ae2366f57d2 · inbound

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards cites this paper.

EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 50

Resolution
unresolved
no resolver link, observed 2026-08-03T21:09:26.512199Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T21:09:26.512199Z digest=sha256:f00f87db87c4d56852256aef3859838212055339c949fed4a748a59c945456bd

Observation e528abfb-ef92-4c2d-b553-e4f06f9e33d9 · inbound

SpatialMosaic: A Multiview VLM Dataset for Partial Visibility cites this paper.

SpatialMosaic: A Multiview VLM Dataset for Partial Visibility MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 47

Resolution
verified exact
arxiv_id, observed 2026-05-16T19:43:20.640781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-05-16T19:42:25.626448Z digest=sha256:23d0d60c694c66a2671351b8bcf08b87b14249089a537ecbcfb0ae245c5f0a5f

Observation cc5a4c69-2201-4625-91c6-b326f0c9c601 · inbound

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy cites this paper.

UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 51

Resolution
unresolved
no resolver link, observed 2026-07-13T18:42:33.178622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-13T18:42:33.178622Z digest=sha256:ceab6b135aa2fb275a14029d2c131cf69358cbc63b1362d4a99a8cd3d32c3184

Observation 98f7788d-89fb-4221-a3a6-8e2fe4349c35 · inbound

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks cites this paper.

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:45:28.188026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T13:41:37.942145Z digest=sha256:7edb6cb7d2e8a74fe61f4241b60cf5d113b6ce9b24f779a31f8132c0f0627d7e

Observation 96528a01-6670-41a2-9ee4-814b23e732ab · inbound

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation cites this paper.

AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 172

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T09:43:49.492561Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-05-10T05:10:44.608959Z digest=sha256:cf068d213cc4b306f046bfe942ac35ee3e300c575454dc2dc7aacbf5d2bb4a3e

Observation 5ac6012a-59d6-48ca-93eb-4dd55a79375c · inbound

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning cites this paper.

Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 199

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T09:07:48.378037Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=arxiv_source observed=2026-06-27T10:28:11.440915Z digest=sha256:12c62eb2f377c8568fd604eccee8431004635d4c4559d985106ebea900ee9908

Observation 8d9b819c-2ffe-4f6f-a328-8344868c714b · inbound

GRIP: Feedback-Guided Prompt Retrieval for Large Multimodal Models cites this paper.

GRIP: Feedback-Guided Prompt Retrieval for Large Multimodal Models MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-07-03T11:28:04.251441Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.

source=pdf_text observed=2026-06-27T09:34:20.560870Z digest=sha256:b9a73947545b1dbe85f632eb0502a0b6e52767591a876a4a2dfff7388c5e8dc2

Observation 0f65fee5-c7bb-441e-af98-09defcb99b27 · inbound

MentalThink: Shaping Thoughts in Mental SVG World cites this paper.

MentalThink: Shaping Thoughts in Mental SVG World MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 184

Resolution
unresolved
no resolver link, observed 2026-07-12T01:50:59.184754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-12T01:50:59.184754Z digest=sha256:e599c13bb6b7a36d83ea6800784142778c42217f39262fc46400e20ba1adeaa1