Pith. sign in

Paper Citation Record · LEDGER

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

As of 7 August 2026, this Paper Citation Record lists 55 of 55 outbound references and 5 inbound Pith citation observations for arXiv:2506.07936.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.07936 v1

Coverage vector

measured 55 of 55 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:25:30.497174Z

measured 60 of 60 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00

measured 5 of 5 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:43:16.986690Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

55 of 55 outbound references displayed

  • verified exact2
  • verified fuzzy11
  • unresolved41
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

0
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation b5eba608-dd3f-43d8-809a-2737621f1a43 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.326469Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.326469Z digest=sha256:e03048bf3e4ebef99b72e981d4a5d2fd010d3cc3489b84da30bbdde4ac6026bf

Observation 6ac019dc-ebe7-4ccd-93b3-2005b827e975 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.330578Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.330578Z digest=sha256:09243655945fa1772c7ba0a0d41b2469f620e6cb54f278afb05acc3330ab638a

Observation 0206e29a-924e-436e-8dc7-d959a0587dfe · outbound

This paper cites Qwen2.5-VL Technical Report.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Qwen2.5-VL Technical Report

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.333944Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.333944Z digest=sha256:bc8102a33ec790df0edde0e0a7c1f03879762a71c73b9fbf5b2239dc5ac31aa1

Observation 66e436cd-5e95-4931-ae50-c3bd46f26429 · outbound

This paper cites What makes multimodal in-context learning work? InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1539–1550, 2024.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models What makes multimodal in-context learning work? InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 1539–1550, 2024

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.337467Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.337467Z digest=sha256:3ab4bf6a25b183f58bef8113e00c58d16627cb8f2749da5b260eb5425e03fa97

Observation 1b60879c-f6ec-499a-a516-da1b4fdfefe6 · outbound

This paper cites Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Language models are few-shot learners.Advances in neural information processing systems, 33:1877–1901, 2020

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.340707Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.340707Z digest=sha256:db0f50d3ac78a1582a0c0f07e74edcf0518952a8ecd2ff3b65a8b570e7974a94

Observation 5cfa213e-4e85-4600-8e5a-d68347522cd7 · outbound

This paper cites M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models M$^3$CoT: A Novel Benchmark for Multi-Domain Multi-step Multi-modal Chain-of-Thought

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.344139Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.344139Z digest=sha256:69778680e8fa769d55ebdb5d45bf299aab1ebb70df3b311c33e15c696a364317

Observation 0a990664-95bc-4ed7-9107-548f6eb72e66 · outbound

This paper cites Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.347857Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.347857Z digest=sha256:0a3cbbdee8b7084d32857b56fe322c59a325b9619993e71959b82f3a8fef7252

Observation d7134849-2ec4-4fe9-9076-ec3d3ab643e4 · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.352098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.352098Z digest=sha256:b49268cb5a4bd5a0d977cd05f872a8c06d00f3f36ca42b8bbedea27a116c02e5

Observation b3be597a-86b0-4857-a4ed-2e67d0c0c170 · outbound

This paper cites A Survey on In-context Learning.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models A Survey on In-context Learning

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.355259Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.355259Z digest=sha256:0c82c7af1217e7d8fae8801549c2cfd3204be3567fc0b2d01431948f15e35669

Observation 3124b6bc-7b85-49ed-8d4a-6cc563b1ffcb · outbound

This paper cites Vlmevalkit: An open-source toolkit for evaluating large multi-modality models.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Vlmevalkit: An open-source toolkit for evaluating large multi-modality models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.358643Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.358643Z digest=sha256:ab077959e1e0b9c47d7c4abbefe723988eb896c66cb4f73fa32cdee62ba8831d

Observation 79eb36fa-ddc5-4e7c-9127-09243a67138f · outbound

This paper cites Interleaved-Modal Chain-of-Thought.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Interleaved-Modal Chain-of-Thought

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.361398Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.361398Z digest=sha256:03cc8f7eb407b9ae454c8a162a9f2a2e37788002b1bd156398fa140f71f5fdb0

Observation caf709a0-2700-4a83-a67d-e1ac8b44295c · outbound

This paper cites Innate Reasoning is Not Enough: In-Context Learning Enhances Reasoning Large Language Models with Less Overthinking.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Innate Reasoning is Not Enough: In-Context Learning Enhances Reasoning Large Language Models with Less Overthinking

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.364635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.364635Z digest=sha256:0c078e4c1f3cff916fe9dd53669e44c8cb2d27d20bfef042ea2ef75e8a04fa1b

Observation 58541b15-d921-4064-b89d-6dee79b1e6b8 · outbound

This paper cites Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Can MLLMs Reason in Multimodality? EMMA: An Enhanced MultiModal ReAsoning Benchmark

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.367946Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.367946Z digest=sha256:05724ca721422f0fa07c04261462ad249c0852fa4e8229bf06c5a824199e47c7

Observation cac6d4c7-64c9-4415-9f84-4ca332bb8d68 · outbound

This paper cites MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models MME-CoT: Benchmarking Chain-of-Thought in Large Multimodal Models for Reasoning Quality, Robustness, and Efficiency

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.371674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.371674Z digest=sha256:a78513872177d9d424baccf270bda6e236690afdf67f101b7b3426d9a7ee7543

Observation 725bc155-8341-46d3-a66d-1a6395d71782 · outbound

This paper cites Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Demonstrate-Search-Predict: Composing retrieval and language models for knowledge-intensive NLP

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.375814Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.375814Z digest=sha256:366155fe356ad811ecfcaa301a96bc857b060ea4fdec1f2bbdb0d8a22e653eca

Observation be9d5f8a-08dc-47bc-a483-72007514d435 · outbound

This paper cites What matters when building vision-language models?.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models What matters when building vision-language models?

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.379020Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.379020Z digest=sha256:c98302047252d37e1c8acf12cba6c9eb106d03a94e06a9cf65600e330a9d59d5

Observation a003c4e3-e160-4134-952d-a5187c342958 · outbound

This paper cites Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Learn to Explain: Multimodal Reasoning via Thought Chains for Science Question Answering

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.381931Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.381931Z digest=sha256:7c19344c944588b634b5dd112b353fc12f1c3d0ec764c409b5230e4f5f527ed8

Observation c9adec86-6858-4838-8f48-a51fcd794183 · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.385062Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.385062Z digest=sha256:ebbd023ebbc0cf187ee285c2f4f1383719b7b28ec2434710930797d12c30eff4

Observation ee0fcb41-a038-4742-940b-4121a9091338 · outbound

This paper cites Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Fantastically Ordered Prompts and Where to Find Them: Overcoming Few-Shot Prompt Order Sensitivity

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.388110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.388110Z digest=sha256:0e4622a963b0e0f7a6be471cecfaa00869481558936389cc220c9cf875c6363c

Observation 6903f7b2-2961-4c17-8107-a935e8baa3e8 · outbound

This paper cites Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.https: //ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ , 2024.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Llama 3.2: Revolutionizing edge ai and vision with open, customizable models.https: //ai.meta.com/blog/llama-3-2-connect-2024-vision-edge-mobile-devices/ , 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:31.049043Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:25:30.391394Z digest=sha256:6ce050adf6ef93313279bf9cd16f406235de727f1f1869dc27bd4d67122a9a7c

Observation 1a023700-31db-4bd3-be62-7c1147e17792 · outbound

This paper cites Few-shot Fine-tuning vs. In-context Learning: A Fair Comparison and Evaluation.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Few-shot Fine-tuning vs. In-context Learning: A Fair Comparison and Evaluation

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.394417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.394417Z digest=sha256:6e1c352e29602cbbb2ad81ffc218ad514b75b14bb749fbd07de5301be8ddc7a1

Observation f8ae2fe6-4446-47d9-9941-f825c3bfb4e9 · outbound

This paper cites Cross-lingual Prompting: Improving Zero-shot Chain-of-Thought Reasoning across Languages.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Cross-lingual Prompting: Improving Zero-shot Chain-of-Thought Reasoning across Languages

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.397479Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.397479Z digest=sha256:a5c23cc52f8afca16bf54bb51bd4817e911708e41e1afebd24c913c1a62fb5f6

Observation 9d60bd3c-8794-4bbd-9635-e66877b7d1e4 · outbound

This paper cites What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.400396Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.400396Z digest=sha256:839a4301852e5e89fb532921c6ee3d1788bda4544c0be5d35798ab39f819e70b

Observation 0d358e1e-a868-4533-be8e-b2a645030683 · outbound

This paper cites an unresolved cited work.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Unresolved cited work

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.403345Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.403345Z digest=sha256:6dc0bf8b352b54bf0ed4e4104d79a67129f6180233764647dd042dfb0470523d

Observation cdd8f034-939e-43fc-bb23-0587c90b662e · outbound

This paper cites A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models A-OKVQA: A Benchmark for Visual Question Answering using World Knowledge

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.406334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.406334Z digest=sha256:6b5346b70164ef239d4a4fa583f80bdfd43c0e86ae59400c9aca5c819c8611d3

Observation 905500d6-2fba-4677-9eb9-ce6a287ce378 · outbound

This paper cites VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models VLM-R1: A Stable and Generalizable R1-style Large Vision-Language Model

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.409388Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.409388Z digest=sha256:1ed1eebdc5713331567b2ba9ed95fe6bc5b3bcf3f03d120160b44717b6c3dc85

Observation c79db2fd-71d1-418f-af04-e7c53c07e56f · outbound

This paper cites Towards vqa models that can read.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Towards vqa models that can read

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:31.040011Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:25:30.412385Z digest=sha256:6d5c4c0dd965ee663b4f1cff61d7c0a49b2d14c4363342a4d7e3f92a399cf385

Observation 51a84144-6660-4ddd-8a42-896548d3a749 · outbound

This paper cites VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.415272Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.415272Z digest=sha256:20d61b8a4ee6cb98ce1b5f78cdd41d876b4295ff194acaa797f04a1ceb78a8f8

Observation 9fef8f3e-b170-4db1-9add-e81f420a39da · outbound

This paper cites Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Enhancing the Reasoning Ability of Multimodal Large Language Models via Mixed Preference Optimization

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.418218Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.418218Z digest=sha256:9fec518a7997f577c364f7fa6e9bcfcc06a2397654e3c9e397dafbc3a56ed799

Observation ed83cb5c-4e55-47a7-a25b-9fdef8c40abb · outbound

This paper cites Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Multimodal Chain-of-Thought Reasoning: A Comprehensive Survey

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.421112Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.421112Z digest=sha256:3a4f55f408390611d207b82c921c6f4ba13164667fbebfbf20545d65eba0f86f

Observation 269f6903-dc82-4e4e-b990-fe6d1046415c · outbound

This paper cites Emergent Abilities of Large Language Models.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Emergent Abilities of Large Language Models

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.424057Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.424057Z digest=sha256:ae78b3562d140cade6197faabeff53716bdec60915e1d2683686b96b65eac21c

Observation 0db51feb-290f-4c93-b4ca-b6f292731f52 · outbound

This paper cites Chain-of-Thought Prompting Elicits Reasoning in Large Language Models.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Chain-of-Thought Prompting Elicits Reasoning in Large Language Models

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.427120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.427120Z digest=sha256:5f23b0f0eb86f14dd73ed33e1c4b49e9df9d69ae330000b9e0d2897fc6747ca6

Observation 0aad04a6-6109-44b6-be93-8f56153d5371 · outbound

This paper cites Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Self-Adaptive In-Context Learning: An Information Compression Perspective for In-Context Example Selection and Ordering

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.430572Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.430572Z digest=sha256:a60f8227543de45ad645db05f7ed97aad994edeb950063470845fa7be16ca309

Observation 736cf377-7b4e-4e2b-8d0a-bb9eff84a94c · outbound

This paper cites Addressing Order Sensitivity of In-Context Demonstration Examples in Causal Language Models.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Addressing Order Sensitivity of In-Context Demonstration Examples in Causal Language Models

Reference 34

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:25:30.646131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:25:30.433663Z digest=sha256:132992e227746a6bebeef84b7035313e8459a768909470e9f65d0e22dc229541

Observation 0a3eec10-ee2b-4186-89bb-5b762e7a11a6 · outbound

This paper cites LLaVA-CoT: Let Vision Language Models Reason Step-by-Step.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models LLaVA-CoT: Let Vision Language Models Reason Step-by-Step

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.436700Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.436700Z digest=sha256:554e6546b86b2ab5978590d94393e05b1abbd01080a70b44f063f9500aabdd35

Observation ef44611e-ed1c-4af7-8c56-8c74eb41f55c · outbound

This paper cites From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models From Introspection to Best Practices: Principled Analysis of Demonstrations in Multimodal In-Context Learning

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.439676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.439676Z digest=sha256:e03937d9d5d905381eb14e0a174ad734e7d78669b8ab92354dd0302c3827a5b0

Observation 56d2d879-eda2-45fa-a7d1-d766326c8e85 · outbound

This paper cites Formal Mathematical Reasoning: A New Frontier in AI.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Formal Mathematical Reasoning: A New Frontier in AI

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.442609Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.442609Z digest=sha256:6ad078def89f259bca54374dff1818541482fbd0473ecc7aedb04de423403747

Observation 35657318-eab6-4585-a89d-41f3582b7ffb · outbound

This paper cites An empirical study of gpt-3 for few-shot knowledge-based vqa.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models An empirical study of gpt-3 for few-shot knowledge-based vqa

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:31.031434Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:25:30.445696Z digest=sha256:baaba1d97a8dc912b791f69b5a2c8a1c4c0fad8aa08cb47c4631e8105f29bc8f

Observation 3d2ce3f7-cb96-49b9-a637-0ebc28f48b68 · outbound

This paper cites Automatic Chain of Thought Prompting in Large Language Models.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Automatic Chain of Thought Prompting in Large Language Models

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.448415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.448415Z digest=sha256:f4f5564484516d41bff279e66423abf70c9f59e6f4a87d061433a145fda4c602

Observation cf7df1ba-8ce8-49cd-91ef-48a135e90ccb · outbound

This paper cites Wong, and Simon See.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Wong, and Simon See

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.451593Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.451593Z digest=sha256:90cff3a4e82ed005fd9ac019398e349b01ac5b50175cc90d7b0374593c28cb49

Observation ae870c12-76c4-4867-99aa-71454747ca65 · outbound

This paper cites Least-to-Most Prompting Enables Complex Reasoning in Large Language Models.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Least-to-Most Prompting Enables Complex Reasoning in Large Language Models

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.454384Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.454384Z digest=sha256:d0b85199ca9502f31660028f07b5d9d383f997f789d9c94d224a211d0b5318fe

Observation 964eca0e-c530-40d4-8056-76e070db338b · outbound

This paper cites VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:30.457496Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:25:30.457496Z digest=sha256:e0c821dd178f6f3218b1942be8bd0a64c61ec71fa868aa44bd2f217f3166da8a

Observation bd652441-f733-40b4-bdb2-c4b0485883a4 · outbound

This paper cites Can In-context Learners Learn a Reasoning Concept from Demonstrations?.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Can In-context Learners Learn a Reasoning Concept from Demonstrations?

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-07T05:25:30.527381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:25:30.460503Z digest=sha256:17f77a3ec3aec4b3a1808081e8913404fcf0d0c0510eca0d8706385734359b31

Observation 4a36758c-10e8-4e96-bdba-3a2e18ee3d68 · outbound

This paper cites Density is calculated as mass/volume.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Density is calculated as mass/volume

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:31.022770Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:25:30.464287Z digest=sha256:7fde236a81e86a281f1c548a06381d319b9764253e0a66d29cab431af4feb422

Observation 574fe031-2fd9-4b20-aa9b-b76a3abdb119 · outbound

This paper cites an unresolved cited work.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Unresolved cited work

Reference 45

Resolution
malformed identifier
raw_fallback, observed 2026-08-07T05:25:31.014289Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:25:30.467233Z digest=sha256:ca893a3f0d7e525734ff21c35d151035c88fd18b0d9324a35007a1b38a28bc51

Observation d43821c7-b720-4112-ab66-d3ff04106176 · outbound

This paper cites an unresolved cited work.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Unresolved cited work

Reference 46

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:25:31.005482Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:25:30.470376Z digest=sha256:80ccc6df394ed2856cc17cf0ce9b0f4f93a180a9a52d99cea2f2bdd60f4e6aa3

Observation 084a3b22-a8ff-479f-9e4c-88a6f6bbcafa · outbound

This paper cites Final answer:.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Final answer:

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:30.996962Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:25:30.473371Z digest=sha256:5674ef3270156d91652e8bc3f1515dee9c3443c2342da2064badf20444595fa6

Observation a5b37fc7-91ce-4e09-860f-d97eaaba693e · outbound

This paper cites In the graphs, we need to compare the export value (top graph) with the import value (bottom graph) for each country in the year 2020.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models In the graphs, we need to compare the export value (top graph) with the import value (bottom graph) for each country in the year 2020

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:30.988293Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:25:30.476297Z digest=sha256:0ab82522f67804f23635fe730c7944a18b56c161ee25f1b2381ed4f0dbb21246

Observation 99a3edf5-38dd-49e6-afcf-e42e0a996b9b · outbound

This paper cites Export > Import, so Country 3 has a surplus.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Export > Import, so Country 3 has a surplus

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:30.978120Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:25:30.479128Z digest=sha256:c2b9bdca1780db055cbba74396a641da97539e0ee0eb9ccfcba420d4413b0046

Observation dcf0a135-2d0a-411b-b19e-a5c855ae7879 · outbound

This paper cites an unresolved cited work.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Unresolved cited work

Reference 50

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:25:30.969273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:25:30.482045Z digest=sha256:715b77132b8f3f0167212b3547b419bda1bfdf1cfd2a1fc9478b29e8d108e8f9

Observation 2a9c27f2-c4f4-41c8-9844-8a492ad6ffd2 · outbound

This paper cites Final answer:.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Final answer:

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:30.959723Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:25:30.484729Z digest=sha256:32adf164b57d0d2c58fcfbb4c90649e97c05c89cc682af1abf63a628cf1a4e69

Observation 4c6fbc48-6bfe-4c30-bba2-6546a303e330 · outbound

This paper cites a² varies inversely with b³.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models a² varies inversely with b³

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:30.949945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:25:30.487895Z digest=sha256:74c58246592e514993e35d6a28203f4532e00517730e75f0539ccdaf6d2b2088

Observation 3e46e54a-673e-47f6-a7ba-0ae533126ade · outbound

This paper cites Therefore, a² = 7² = 49.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Therefore, a² = 7² = 49

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:30.941221Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:25:30.491118Z digest=sha256:4336ebb3942a232586dbdd9ff90458e43bfd8182b638de136e376b17ad008d6f

Observation e64c6f14-4372-42e7-b695-b4ca36a8aab3 · outbound

This paper cites When b = 6, a² = 1323 / 6³ = 1323 / 216 = 6.125.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models When b = 6, a² = 1323 / 6³ = 1323 / 216 = 6.125

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:30.932322Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:25:30.494019Z digest=sha256:9c7eeb784ce46e0d9e4dabbb0bb0c9f2133e297040807b6c18d76e8ab1e96f81

Observation 54bdac12-e0f8-42b8-92ef-2625aef72dec · outbound

This paper cites an unresolved cited work.

Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models Unresolved cited work

Reference 55

Resolution
unresolved
raw_fallback, observed 2026-08-07T05:25:30.923432Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-08-07T05:25:30.497174Z digest=sha256:f4c0b5afea3112366e1009a4accd496ee1c2e0ba33fb7c977122f349ef2ce55f

Pith citing papers

Observation adec5307-bff3-4d51-9fa7-21126ac3a9bb · inbound

In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems cites this paper.

In-context Learning of Vision Language Models for Detection of Physical and Digital Attacks against Face Recognition Systems Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

Reference 90

Resolution
unresolved
no resolver link, observed 2026-08-06T15:43:16.986690Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:43:16.986690Z digest=sha256:1c8fd0920fde9cd87e0c97f66df2404de2e4ad7eab0575dce45a2497b4519eef

Observation f5ca7f5c-4333-4f1f-ac79-bab58917cdb4 · inbound

True Multimodal In-Context Learning Needs Attention to the Visual Context cites this paper.

True Multimodal In-Context Learning Needs Attention to the Visual Context Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.173300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.173300Z digest=sha256:d2cb1729c6c0f8a98df4dea9612c17f4a251a6831da4ed77a13db5b398144b4d

Observation 84d6614a-312c-4db4-b403-35f50e49d566 · inbound

MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction cites this paper.

MetaEmbed: Scaling Multimodal Retrieval at Test-Time with Flexible Late Interaction Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

Reference 23

Resolution
verified exact
arxiv_id, observed 2026-05-18T14:11:27.448460Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-18T14:09:22.942238Z digest=sha256:ab8a7b5fe249aeb2119487a8f4ced638b821a5bf6d04f85a2507529dbdf089ec

Observation 957de704-b9a9-4c56-983f-e25012054cec · inbound

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks cites this paper.

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

Reference 16

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:45:28.245354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=arxiv_source observed=2026-05-10T13:41:37.942145Z digest=sha256:524249f0407da602b0d313963c873978adbc09e7e8401e59c417633051f22780

Observation 1dae7721-237e-4948-8945-f6b27325da46 · inbound

OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice cites this paper.

OralMLLM-Bench: Evaluating Cognitive Capabilities of Multimodal Large Language Models in Dental Practice Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

Reference 44

Resolution
verified exact
arxiv_id, observed 2026-05-11T01:05:50.173877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.

source=pdf_text observed=2026-05-11T01:03:50.055081Z digest=sha256:5d11b72962c537d562aaebf634c486bff0f38fbe52082e16dd18aefe1e2cf7b2