Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:25:30.497174Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 5 inbound Pith citation observations for arXiv:2506.07936.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T05:25:30.497174Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T15:43:16.986690Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z
55 of 55 outbound references displayed
External citation measurements
0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z
Observation b5eba608-dd3f-43d8-809a-2737621f1a43 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ac019dc-ebe7-4ccd-93b3-2005b827e975 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0206e29a-924e-436e-8dc7-d959a0587dfe · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Qwen2.5-VL Technical Report
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66e436cd-5e95-4931-ae50-c3bd46f26429 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models What makes multimodal in-context learning work? InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1539–1550, 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b60879c-f6ec-499a-a516-da1b4fdfefe6 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cfa213e-4e85-4600-8e5a-d68347522cd7 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0a990664-95bc-4ed7-9107-548f6eb72e66 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d7134849-2ec4-4fe9-9076-ec3d3ab643e4 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b3be597a-86b0-4857-a4ed-2e67d0c0c170 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models A Survey on In-context Learning
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3124b6bc-7b85-49ed-8d4a-6cc563b1ffcb · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Vlmevalkit: An open-source toolkit for evaluating large multi-modality models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 79eb36fa-ddc5-4e7c-9127-09243a67138f · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Interleaved-Modal Chain-of-Thought
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation caf709a0-2700-4a83-a67d-e1ac8b44295c · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Innate Reasoning is Not Enough: In-Context Learning Enhances Reasoning Large Language Models with Less Overthinking
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 58541b15-d921-4064-b89d-6dee79b1e6b8 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cac6d4c7-64c9-4415-9f84-4ca332bb8d68 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 725bc155-8341-46d3-a66d-1a6395d71782 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be9d5f8a-08dc-47bc-a483-72007514d435 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models What matters when building vision-language models?
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a003c4e3-e160-4134-952d-a5187c342958 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9adec86-6858-4838-8f48-a51fcd794183 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ee0fcb41-a038-4742-940b-4121a9091338 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6903f7b2-2961-4c17-8107-a935e8baa3e8 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.https: //ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ , 2024
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1a023700-31db-4bd3-be62-7c1147e17792 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Few-shot Fine-tuning vs. In-context Learning: A Fair Comparison and Evaluation
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f8ae2fe6-4446-47d9-9941-f825c3bfb4e9 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Cross-lingual Prompting: Improving Zero-shot Chain-of-Thought Reasoning across Languages
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d60bd3c-8794-4bbd-9635-e66877b7d1e4 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0d358e1e-a868-4533-be8e-b2a645030683 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Unresolved cited work
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cdd8f034-939e-43fc-bb23-0587c90b662e · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 905500d6-2fba-4677-9eb9-ce6a287ce378 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c79db2fd-71d1-418f-af04-e7c53c07e56f · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Towards vqa models that can read
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 51a84144-6660-4ddd-8a42-896548d3a749 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9fef8f3e-b170-4db1-9add-e81f420a39da · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ed83cb5c-4e55-47a7-a25b-9fdef8c40abb · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 269f6903-dc82-4e4e-b990-fe6d1046415c · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Emergent Abilities of Large Language Models
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0db51feb-290f-4c93-b4ca-b6f292731f52 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aad04a6-6109-44b6-be93-8f56153d5371 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 736cf377-7b4e-4e2b-8d0a-bb9eff84a94c · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Addressing Order Sensitivity of In-Context Demonstration Examples in Causal Language Models
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0a3eec10-ee2b-4186-89bb-5b762e7a11a6 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models LLaVA-CoT: Let Vision Language Models Reason Step-by-Step
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef44611e-ed1c-4af7-8c56-8c74eb41f55c · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56d2d879-eda2-45fa-a7d1-d766326c8e85 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Formal Mathematical Reasoning: A New Frontier in AI
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 35657318-eab6-4585-a89d-41f3582b7ffb · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models An empirical study of gpt-3 for few-shot knowledge-based vqa
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3d2ce3f7-cb96-49b9-a637-0ebc28f48b68 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Automatic Chain of Thought Prompting in Large Language Models
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf7df1ba-8ce8-49cd-91ef-48a135e90ccb · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Wong, and Simon See
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae870c12-76c4-4867-99aa-71454747ca65 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Least-to-Most Prompting Enables Complex Reasoning in Large Language Models
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 964eca0e-c530-40d4-8056-76e070db338b · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd652441-f733-40b4-bdb2-c4b0485883a4 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Can In-context Learners Learn a Reasoning Concept from Demonstrations?
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4a36758c-10e8-4e96-bdba-3a2e18ee3d68 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Density is calculated as mass/volume
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 574fe031-2fd9-4b20-aa9b-b76a3abdb119 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation d43821c7-b720-4112-ab66-d3ff04106176 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 084a3b22-a8ff-479f-9e4c-88a6f6bbcafa · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Final answer:
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation a5b37fc7-91ce-4e09-860f-d97eaaba693e · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models In the graphs, we need to compare the export value (top graph) with the import value (bottom graph) for each country in the year 2020
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 99a3edf5-38dd-49e6-afcf-e42e0a996b9b · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Export > Import, so Country 3 has a surplus
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation dcf0a135-2d0a-411b-b19e-a5c855ae7879 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 2a9c27f2-c4f4-41c8-9844-8a492ad6ffd2 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Final answer:
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4c6fbc48-6bfe-4c30-bba2-6546a303e330 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models a² varies inversely with b³
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 3e46e54a-673e-47f6-a7ba-0ae533126ade · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Therefore, a² = 7² = 49
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation e64c6f14-4372-42e7-b695-b4ca36a8aab3 · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models When b = 6, a² = 1323 / 6³ = 1323 / 216 = 6.125
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 54bdac12-e0f8-42b8-92ef-2625aef72dec · outbound
Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Unresolved cited work
Reference 55
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation adec5307-bff3-4d51-9fa7-21126ac3a9bb · inbound
In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models
Reference 90
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f5ca7f5c-4333-4f1f-ac79-bab58917cdb4 · inbound
True Multimodal In-Context Learning Needs Attention to the Visual Context Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models
Reference 2021
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 84d6614a-312c-4db4-b403-35f50e49d566 · inbound
MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 957de704-b9a9-4c56-983f-e25012054cec · inbound
Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1dae7721-237e-4948-8945-f6b27325da46 · inbound
OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.