Pith. sign in

Paper Citation Record · LEDGER

True Multimodal In-Context Learning Needs Attention to the Visual Context

As of 17 August 2026, this Paper Citation Record lists 33 of 33 outbound references and 4 inbound Pith citation observations for arXiv:2507.15807.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.15807 v2

Coverage vector

measured 33 of 33 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T15:29:35.934009Z

measured 37 of 37 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 4 of 4 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T07:15:11.031811Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-02T20:47:22.870706Z

Reference resolution

33 of 33 outbound references displayed

  • verified exact2
  • verified fuzzy4
  • unresolved27
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation e30b601e-a2b9-431a-aed0-d4e28defce69 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

True Multimodal In-Context Learning Needs Attention to the Visual Context Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:33.466845Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:33.466845Z digest=sha256:0b2a807fe393eed852abf8c3aefce7fb8f01376bf45ef2fce0856f8b40151b31

Observation e4777390-faa6-4602-84eb-a7e7ce3d8d9f · outbound

This paper cites Each task is designed with adjustable difficulty levels, such as more diverse con- cepts in novel concept binding, more complex visual patterns in pattern interpretation, etc.

True Multimodal In-Context Learning Needs Attention to the Visual Context Each task is designed with adjustable difficulty levels, such as more diverse con- cepts in novel concept binding, more complex visual patterns in pattern interpretation, etc

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:29:36.680644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:29:35.671431Z digest=sha256:e695b8e461ea6ebb7d6d14ba8fa3b39fdbc8daeffb8cb6a9f90c838488db888a

Observation a5ea653a-54bd-4e0a-b88c-e7c1f2c6168f · outbound

This paper cites A Survey on In-context Learning.

True Multimodal In-Context Learning Needs Attention to the Visual Context A Survey on In-context Learning

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:33.727729Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:33.727729Z digest=sha256:8f2aa89321d72aa97d36cd48cc1561b08669896e234ae453e7c5eeeeadd56c43

Observation efe63c4b-2f37-4df8-80bc-62a459a46624 · outbound

This paper cites In-context learning enables multimodal large language models to classify cancer pathology images.

True Multimodal In-Context Learning Needs Attention to the Visual Context In-context learning enables multimodal large language models to classify cancer pathology images

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:33.873679Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:33.873679Z digest=sha256:8ee45fd7489fb04d367c98ee4428c5dc69a6fad660b1fac47ac108f60b09ba2f

Observation 7e94f6ba-10e7-4fbf-a878-8bbf6543c081 · outbound

This paper cites In-context learning enables multimodal large language models to classify cancer pathology images.

True Multimodal In-Context Learning Needs Attention to the Visual Context In-context learning enables multimodal large language models to classify cancer pathology images

Reference 8

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:29:36.196407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:29:33.947604Z digest=sha256:87ff95abfa4cf529f3ecaf7bc9e2712dff5067f2218d2329621493b9f1aaa1f1

Observation 81541ab2-1b20-4b79-8230-ef2d6bf3d74c · outbound

This paper cites Many-Shot In-Context Learning in Multimodal Foundation Models.

True Multimodal In-Context Learning Needs Attention to the Visual Context Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.320527Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.320527Z digest=sha256:1d2d36e4001a912c0cc1d837c6e87dddceb70acc3e16a32019f641cff081e65a

Observation f949f0e5-46a8-445c-be4f-4a7c575cc8a6 · outbound

This paper cites Many-Shot In-Context Learning in Multimodal Foundation Models.

True Multimodal In-Context Learning Needs Attention to the Visual Context Many-Shot In-Context Learning in Multimodal Foundation Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.377976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.377976Z digest=sha256:4a33337676ca422ebc7f2b6b2f718f93aa5d2ebbc9c6b53d5a907306c5ff1328

Observation 3148e369-a2de-4a8f-aab3-1ace73c11c2d · outbound

This paper cites MIBench: Evaluating Multimodal Large Language Models over Multiple Images.

True Multimodal In-Context Learning Needs Attention to the Visual Context MIBench: Evaluating Multimodal Large Language Models over Multiple Images

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.466045Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.466045Z digest=sha256:41d434ba54b6fdf6015ab4ae21d32118ecc99ffe59b6097330ba1a9d53863523

Observation dc9c221e-66ce-4533-8436-62b57fed643d · outbound

This paper cites A survey on lora of large language models.

True Multimodal In-Context Learning Needs Attention to the Visual Context A survey on lora of large language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:29:36.710394Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:29:34.538808Z digest=sha256:bcbaa7af5b52c8b8ce32a6b032adc9996ad384841a2c677a569c25e3d9ecf530

Observation 97e71c88-e3a7-4d53-8d52-d14e98fc31ea · outbound

This paper cites GPT-4o System Card.

True Multimodal In-Context Learning Needs Attention to the Visual Context GPT-4o System Card

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.609047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.609047Z digest=sha256:e5df6745c43e3c399c782edc8463a22f5c38a3004ae8017911709f473bf2ce49

Observation c9e5bc96-48c3-4d02-83a9-7a99df77082b · outbound

This paper cites GPT-4o System Card.

True Multimodal In-Context Learning Needs Attention to the Visual Context GPT-4o System Card

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.684005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.684005Z digest=sha256:f3f3ff9ac4904f16c7045af315193d3655a5ed24b7a47d7bc182524e9c9f5715

Observation 9f34ef8a-b5f5-4b23-88f1-b4bde65a90f3 · outbound

This paper cites What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration.

True Multimodal In-Context Learning Needs Attention to the Visual Context What Factors Affect Multi-Modal In-Context Learning? An In-Depth Exploration

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.782618Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.782618Z digest=sha256:9882cf510240883585ad42247dc941b648474f9995b36db794e52081657d676c

Observation 4bb3b7ee-5dff-4e22-abba-519bf33b5195 · outbound

This paper cites A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications.

True Multimodal In-Context Learning Needs Attention to the Visual Context A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.853072Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.853072Z digest=sha256:dea62b1d5cd24b4dd109dbd0d7f0afd9fad9906720fd95c74e5e13fe463a1c1d

Observation cf365a9e-c200-4616-a198-f5ed12c00cd5 · outbound

This paper cites A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications.

True Multimodal In-Context Learning Needs Attention to the Visual Context A Systematic Survey of Prompt Engineering in Large Language Models: Techniques and Applications

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.925120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.925120Z digest=sha256:9491a078de96616b9a895cd4097d257dd496c8a053af8f8825f109159870c4c9

Observation 7128cb3a-d652-462f-9b52-9734428ecd8b · outbound

This paper cites Learning to Retrieve In-Context Examples for Large Language Models.

True Multimodal In-Context Learning Needs Attention to the Visual Context Learning to Retrieve In-Context Examples for Large Language Models

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:35.040526Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:35.040526Z digest=sha256:099a9600be37d96692856f3094649cf7381fbc56ff59ff87516a8d778d14f6c6

Observation 535cdd76-764d-4c93-90c2-c5102b1a572a · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

True Multimodal In-Context Learning Needs Attention to the Visual Context Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:35.096401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:35.096401Z digest=sha256:8986e139c360f82d8ca0b4571007b8a0e907b182d5a345b8470443e86d889d70

Observation 8d74c232-7be3-496e-8ba4-d34d5a0e7702 · outbound

This paper cites Low-rank adaptation for foundation models: A comprehensive review.

True Multimodal In-Context Learning Needs Attention to the Visual Context Low-rank adaptation for foundation models: A comprehensive review

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:35.167322Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:35.167322Z digest=sha256:8e1705fee5927d2e46fc0ab440cba434430f0b04d4d9b60e00d024535c13d151

Observation 6cec880b-4fb4-4243-b9f4-8972be25ee3b · outbound

This paper cites In-Context Example Selection via Similarity Search Improves Low-Resource Machine Translation.

True Multimodal In-Context Learning Needs Attention to the Visual Context In-Context Example Selection via Similarity Search Improves Low-Resource Machine Translation

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:35.243089Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:35.243089Z digest=sha256:1de725a2f9b7531e85bb79a2dd43ee88d104c2ae552d65bcf66b7df11b102a16

Observation cf00b6b2-469d-4226-94e6-843f3fb610e1 · outbound

This paper cites In-Context Example Selection via Similarity Search Improves Low-Resource Machine Translation.

True Multimodal In-Context Learning Needs Attention to the Visual Context In-Context Example Selection via Similarity Search Improves Low-Resource Machine Translation

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:35.291674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:35.291674Z digest=sha256:6cd62260518a4cd9690f45d1e64d3e47ffc9fd702126a674356e7002f1aa3bab

Observation f304a5c6-2cb3-4968-9053-1d34374fe6bb · outbound

This paper cites MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning.

True Multimodal In-Context Learning Needs Attention to the Visual Context MMICL: Empowering Vision-language Model with Multi-Modal In-Context Learning

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:35.387308Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:35.387308Z digest=sha256:4d93fb3c23ce27bfc1fbf1504dd82f85a1926ee2c4de37efdc741c7fca2613da

Observation 993d4abe-3f16-40b6-9fbd-24509c847c6e · outbound

This paper cites VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning.

True Multimodal In-Context Learning Needs Attention to the Visual Context VL-ICL Bench: The Devil in the Details of Multimodal In-Context Learning

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:35.474334Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:35.474334Z digest=sha256:2a16050b5eef5df41d7e2bf06e83cbe92a11a90eaa7088ef1282c38a564575fb

Observation 72745e32-46c3-46c6-a753-4f3062bbbf79 · outbound

This paper cites Suppose we introduce an attention reallocation factor to the softmax operation by defining F := diag(f) ∈ RL×L, where f ∈ RL is a vector of learnable factors.

True Multimodal In-Context Learning Needs Attention to the Visual Context Suppose we introduce an attention reallocation factor to the softmax operation by defining F := diag(f) ∈ RL×L, where f ∈ RL is a vector of learnable factors

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:29:36.695880Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:29:35.565985Z digest=sha256:bc1d7b7bbe7c47a2f522acacdebba2504bf765615003ea33ab6aafaff3cbf849

Observation 76b1a58c-90ab-4a17-9672-254d8e38bfed · outbound

This paper cites an unresolved cited work.

True Multimodal In-Context Learning Needs Attention to the Visual Context Unresolved cited work

Reference 32

Resolution
unresolved
raw_fallback, observed 2026-08-06T15:29:36.665811Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:29:35.793693Z digest=sha256:8d21bdee696fa942742e4e3ed05041bb217c07e6d105160aa81e28715b56997c

Observation bee96674-cf52-4d2f-ad4a-a0b00a25105c · outbound

This paper cites As shown in the figure, within the range of a few hundred parameters, different configurations have no significant difference in the impact on the final performance.

True Multimodal In-Context Learning Needs Attention to the Visual Context As shown in the figure, within the range of a few hundred parameters, different configurations have no significant difference in the impact on the final performance

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T15:29:36.652176Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:29:35.934009Z digest=sha256:2bcca729ec9697bbfa1c75f9288c960b4ab56b5c34e9af2bb9d3c22d40097f7a

Observation 97a50230-23ed-4cd0-a526-df69a1fde58d · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

True Multimodal In-Context Learning Needs Attention to the Visual Context OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 2015

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:33.523551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:33.523551Z digest=sha256:905b8ffcae4b50b7d9b74ce357485ccb050ccccde45691e01ab112c75ce5ff9d

Observation 84e408d6-aba9-4656-aff9-84ee0bb82e34 · outbound

This paper cites Advanced Multimodal Deep Learning Architecture for Image-Text Matching.

True Multimodal In-Context Learning Needs Attention to the Visual Context Advanced Multimodal Deep Learning Architecture for Image-Text Matching

Reference 2017

Resolution
verified exact
local_arxiv, observed 2026-08-06T15:29:36.122777Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-06T15:29:34.962587Z digest=sha256:9f35eacc1733ab7b88ac23524a87b2fda818b515c303dd2517d6df3b4c9a50c1

Observation 712a4f79-f2dc-43e6-ab4e-2a170e58783a · outbound

This paper cites SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization.

True Multimodal In-Context Learning Needs Attention to the Visual Context SymDPO: Boosting In-Context Learning of Large Multimodal Models with Symbol Demonstration Direct Preference Optimization

Reference 2019

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.240882Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.240882Z digest=sha256:4a85d4dd750f38818e15deae3685ba581fb337a379018deb71f6ed106fb32def

Observation d0efde7c-62a4-4f48-934c-376553f6b27a · outbound

This paper cites Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?.

True Multimodal In-Context Learning Needs Attention to the Visual Context Can Multimodal Large Language Models Truly Perform Multimodal In-Context Learning?

Reference 2020

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:33.637147Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:33.637147Z digest=sha256:b77f4c7d87ffb2bbacf471418eb33d89fa46f20cfac7ad8961f9b9734107c174

Observation f5ca7f5c-4333-4f1f-ac79-bab58917cdb4 · outbound

This paper cites Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models.

True Multimodal In-Context Learning Needs Attention to the Visual Context Mimicking or Reasoning: Rethinking Multi-Modal In-Context Learning in Vision-Language Models

Reference 2021

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.173300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.173300Z digest=sha256:d78786498802345ef088b86daeb41f24eb294df48ea7f7e3b7579f00c211045a

Observation da1fab29-7ff0-4d7d-800c-e27fe95aa35b · outbound

This paper cites Towards Multimodal In-Context Learning for Vision & Language Models.

True Multimodal In-Context Learning Needs Attention to the Visual Context Towards Multimodal In-Context Learning for Vision & Language Models

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:33.800157Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:33.800157Z digest=sha256:996740af915701810c2dc9924788a0588f5f7586988a2c02b89f2954b348a6b3

Observation 6621e578-bf53-499a-ad72-264555676780 · outbound

This paper cites LoRA: Low-Rank Adaptation of Large Language Models.

True Multimodal In-Context Learning Needs Attention to the Visual Context LoRA: Low-Rank Adaptation of Large Language Models

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.121756Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.121756Z digest=sha256:861a9c0a4b6283110943d097368c81f6bc63801333754ec096528c9f4bcb2602

Observation 95b3e3fd-2f3f-4d75-a503-2094b83ceb13 · outbound

This paper cites Language models are few-shot learners.

True Multimodal In-Context Learning Needs Attention to the Visual Context Language models are few-shot learners

Reference 2024

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:33.589687Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:33.589687Z digest=sha256:468d8995121cc203639549578c5e4eeb5efa739c488988239933047562bb386e

Observation 56635091-ef64-4704-98cf-59e530b62c46 · outbound

This paper cites A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models.

True Multimodal In-Context Learning Needs Attention to the Visual Context A Systematic Survey of Prompt Engineering on Vision-Language Foundation Models

Reference 2025

Resolution
unresolved
no resolver link, observed 2026-08-06T15:29:34.010676Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T15:29:34.010676Z digest=sha256:3c5597de0a236f93ca7d4e6ad5e2d13c8364e26a88411a808914db5986ab2f43

Pith citing papers

Observation 228ec975-998a-422c-89b7-0bb4354606c1 · inbound

Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers cites this paper.

Dissecting Multimodal In-Context Learning: Modality Asymmetries and Circuit Dynamics in modern Transformers True Multimodal In-Context Learning Needs Attention to the Visual Context

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-03T07:15:11.031811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T07:15:11.031811Z digest=sha256:e86e083e40069ac2cdd7b42539e0b87b1c0ed963c2fe82fe490ffab488c1b6fd

Observation 49fda585-c582-4230-9647-634ebf07c115 · inbound

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks cites this paper.

Why Multimodal In-Context Learning Lags Behind? Unveiling the Inner Mechanisms and Bottlenecks True Multimodal In-Context Learning Needs Attention to the Visual Context

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:45:28.210983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-05-10T13:41:37.942145Z digest=sha256:ac2b92de8cfa91ef22e8eaee4278a91fc89511d8f989cb4cf3517eed33d28c9c

Observation 694c1bcf-b7dd-4af0-bf36-3dc8c2ea6832 · inbound

Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning cites this paper.

Enhancing Multimodal In-Context Learning via Inductive-Deductive Reasoning True Multimodal In-Context Learning Needs Attention to the Visual Context

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-09T05:55:31.182673Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-05-08T19:21:30.235583Z digest=sha256:fe96319f31e5fc3409d3d83048969b78cdbe11e6b2c8eec9f1e3ccfd9f0efafd

Observation f3145f01-dd25-4afc-8ca7-d6629c2f08fe · inbound

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems cites this paper.

Sci-Rho: A Multilingual Visually-Grounded Symbolic Benchmark for STEM Problems True Multimodal In-Context Learning Needs Attention to the Visual Context

Reference 50

Resolution
verified exact
arxiv_id, observed 2026-07-02T20:47:22.872136Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=arxiv_source observed=2026-06-27T20:11:02.445626Z digest=sha256:e7a022d224442e01daf67e1745f259b8d06e1fbaa086d424466ff6de3b3ee747