Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
Paper Citation Record · LEDGER
As of 11 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 25 inbound Pith citation observations for arXiv:2309.07915.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-11T06:34:44.6726+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-10T22:56:54.862312Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-07-03T11:28:04.249589Z
0 of 0 outbound references displayed
External citation measurements
No source-named external measurement is stored.
No outbound reference observations are available for this paper version.
Observation c003e649-d936-4595-ac39-8a0569c65b9f · inbound
Otter: A Multi-Modal Model with In-Context Instruction Tuning MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 104
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 49c71f5b-2ca1-4f26-ba5c-ee4fda747cc5 · inbound
MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 61
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation c32ec879-219d-46c8-a17e-5888f035abde · inbound
A Survey on Multimodal Large Language Models MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 178
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8f8cd8e6-452f-4c3f-a7ac-9da087753530 · inbound
InternLM-XComposer: A Vision-Language Large Model for Advanced Text-image Comprehension and Composition MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 102
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 7b199355-444a-4547-a8df-919470dc953e · inbound
SPHINX: The Joint Mixing of Weights, Tasks, and Visual Embeddings for Multi-modal Large Language Models MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation dff90377-1243-4301-8102-07998a1a663e · inbound
MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 93
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 82d29ec2-bf0b-4894-a809-45bbdbbc10e9 · inbound
AppAgent: Multimodal Agents as Smartphone Users MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 63612e85-d06c-44c0-85a0-51c889966839 · inbound
Hallucination of Multimodal Large Language Models: A Survey MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 215
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 873119cc-23fb-4666-8c34-0fbe0df37024 · inbound
mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 270
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 6e32069a-5927-46d7-87ed-ff66f9af7781 · inbound
MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 70
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5d7c42c9-d3fb-4109-9bfc-10a9d1ece2e5 · inbound
Embodied VideoAgent: Persistent Memory from Egocentric Videos and Embodied Sensors Enables Dynamic Scene Understanding MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 62
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 980f9d5a-59ed-459f-943b-52df29e9c27d · inbound
Efficiently Enhancing General Agents With Hierarchical-categorical Memory MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7cadc4d-47b2-4570-921e-e97df1d5db9c · inbound
LLaVA-SP: Enhancing Visual Representation with Visual Spatial Tokens for MLLMs MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 65
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 17b80f1f-639d-4ed0-8004-9db7e39d7fca · inbound
DisCo: Towards Distinct and Coherent Visual Encapsulation in Video MLLMs MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 84
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 26378829-0ef0-4a8b-8169-bac28bf660ad · inbound
Beyond Task-Specific Reasoning: A Unified Conditional Generative Framework for Abstract Visual Reasoning MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 61
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f304a5c6-2cb3-4968-9053-1d34374fe6bb · inbound
True Multimodal In-Context Learning Needs Attention to the Visual Context MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc0acddf-0452-4542-8916-34dbb5c29744 · inbound
Region-Level Context-Aware Multimodal Understanding MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c849f00e-97e6-47c5-bef1-9ae2366f57d2 · inbound
EvoLMM: Self-Evolving Large Multimodal Models with Continuous Rewards MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e528abfb-ef92-4c2d-b553-e4f06f9e33d9 · inbound
SpatialMosaic: A Multiview VLM Dataset for Partial Visibility MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation cc5a4c69-2201-4625-91c6-b326f0c9c601 · inbound
UniICL: Systematizing Unified Multimodal In-context Learning through a Capability-Oriented Taxonomy MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98f7788d-89fb-4221-a3a6-8e2fe4349c35 · inbound
Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 96528a01-6670-41a2-9ee4-814b23e732ab · inbound
AnchorSeg: Language Grounded Query Banks for Reasoning Segmentation MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 172
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 5ac6012a-59d6-48ca-93eb-4dd55a79375c · inbound
Task-Aware Structured Memory for Dynamic Multi-modal In-Context Learning MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 199
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 8d9b819c-2ffe-4f6f-a328-8344868c714b · inbound
GRIP: Feedback-Guided Prompt Retrieval for Large Multimodal Models MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-11T06:34:44.6726+00:00.
Observation 0f65fee5-c7bb-441e-af98-09defcb99b27 · inbound
MentalThink: Shaping Thoughts in Mental SVG World MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning
Reference 184
Source-reported events for the cited work
Unavailable: canonical work link unavailable.