Pith. sign in

Paper Citation Record · LEDGER

Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

As of 9 August 2026, this Paper Citation Record lists 0 of 0 outbound references and 33 inbound Pith citation observations for arXiv:2401.06209.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2401.06209 v2

Coverage vector

measured 0 of 0 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links

measured 33 of 33 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-09T06:31:02.800959+00:00

measured 33 of 33 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-07T14:45:32.061891Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Reference resolution

0 of 0 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved0
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

7
arxiv_reference, observed 2026-08-05T02:28:24.338817Z

Outbound references

No outbound reference observations are available for this paper version.

Pith citing papers

Observation acac3d1d-5e78-403e-9081-b9765ca49f05 · inbound

DeepSeek-VL: Towards Real-World Vision-Language Understanding cites this paper.

DeepSeek-VL: Towards Real-World Vision-Language Understanding Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-11T17:58:54.709008Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T17:58:54.177359Z digest=sha256:3cc691140463986c7a2a0d41d6bf940e47ce0c10ee2817f5a10e6fa7a22123c3

Observation 939e1045-4e5a-4cc2-8d49-572279dcc5e2 · inbound

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training cites this paper.

MM1: Methods, Analysis & Insights from Multimodal LLM Pre-training Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 108

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T04:09:36.274795Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T04:09:36.019146Z digest=sha256:34169ed6d3672e3073ab9b13731e4fcdda2b92016215394f1b1a7571a73026b8

Observation 62c29b44-7eb0-47e4-91b2-6be587e15c2a · inbound

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites cites this paper.

How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 111

Resolution
verified exact
arxiv_id, observed 2026-05-12T20:58:59.207454Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-12T20:58:58.849040Z digest=sha256:cc77126ca5352513063d901bbe3fc253eee6740ac16e467c0f82bad829629d51

Observation 6c02f745-d645-4819-9829-623ccc82d368 · inbound

Hallucination of Multimodal Large Language Models: A Survey cites this paper.

Hallucination of Multimodal Large Language Models: A Survey Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 156

Resolution
verified exact
arxiv_id, observed 2026-05-11T12:33:33.027756Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-11T12:33:32.631346Z digest=sha256:f9d53c364e715b83a8889dda2dd2cbbdbd467162e9b4d95825a81117993ef65c

Observation e4e60157-5b8f-4138-bf26-7e5b1f2d07bd · inbound

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations cites this paper.

Seeing It or Not? Interpretable Vision-aware Latent Steering to Mitigate Object Hallucinations Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 65

Resolution
malformed identifier
no resolver link, observed 2026-08-07T14:45:32.061891Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T14:45:32.061891Z digest=sha256:c1e43a9407771666438d754820f179351b361db9dd981bbddd37f37c1ac52060

Observation 91e87091-4d6a-46ae-b4f2-9f348ae857a9 · inbound

AMIA: Automatic Masking and Joint Intention Analysis Makes LVLMs Robust Jailbreak Defenders cites this paper.

AMIA: Automatic Masking and Joint Intention Analysis Makes LVLMs Robust Jailbreak Defenders Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T12:35:28.325476Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T12:35:28.325476Z digest=sha256:158a516fe5e1f47eb307909a564d152ae1cf3f6ccc9a19b557eac67b42141440

Observation 514944c9-b9b9-427f-99fb-7d2e0502644a · inbound

Humanoid World Models: Open World Foundation Models for Humanoid Robotics cites this paper.

Humanoid World Models: Open World Foundation Models for Humanoid Robotics Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T11:54:57.006847Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:54:57.006847Z digest=sha256:54551daf85343ffe9ddd1421d614cea98bae6ba709a3d4f8138059433348dbbe

Observation bafe4089-8d7e-40e5-aae8-f71239d64ec9 · inbound

MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs cites this paper.

MoDA: Modulation Adapter for Fine-Grained Visual Grounding in Instructional MLLMs Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T11:39:46.653047Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T11:39:46.653047Z digest=sha256:daddf27d5ce11d43d6968f5d228536a450e07d28b8ff70f07c517cb96ed16ba1

Observation c4d2c4f5-7971-4722-9fd1-74b1712a1143 · inbound

Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations cites this paper.

Unfolding Spatial Cognition: Evaluating Multimodal Models on Visual Simulations Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T10:41:23.165468Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T10:41:23.165468Z digest=sha256:913183242e9f7d1250985e9905d752dbb303d7dbe7505068f726bf25c0f8d899

Observation 03aea609-f8a5-465a-bd29-24afa8a79559 · inbound

Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation cites this paper.

Backdoor Attack on Vision Language Models with Stealthy Semantic Manipulation Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-07T05:45:56.153036Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:45:56.153036Z digest=sha256:67ca1fde25c6e629d8f9186556f42b97e1a032a162f45f840af348484037fd4e

Observation a3698515-14c1-45aa-a30f-59ac2fb2f97a · inbound

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks cites this paper.

A Good CREPE needs more than just Sugar: Investigating Biases in Compositional Vision-Language Benchmarks Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:21:35.113327Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T05:21:35.113327Z digest=sha256:c3b4e4d949a4c2ecbd6540996494973e72d9ab58f71bfaeafec74963c8fc5735

Observation ffedd751-f1f1-40a0-92bd-419d0354ae2e · inbound

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks cites this paper.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 58

Resolution
verified exact
arxiv_id, observed 2026-05-19T05:57:08.023724Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:a540c3be74bf78349e5d475346af5e88aa8907f2f430a14e319c536f30177327

Observation fa75da25-bfe4-4e13-822e-b46c35599431 · inbound

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model cites this paper.

LLaVA-Critic-R1: Your Critic Model is Secretly a Strong Policy Model Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 43

Resolution
unresolved
no resolver link, observed 2026-08-05T13:24:39.817974Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-05T13:24:39.817974Z digest=sha256:d5ec72834e6e33b3b5907f97de60c3ed944e3a365e0ea25e38d14c46741a2f9b

Observation 1480ea88-014d-41e3-91a9-1e552c22266b · inbound

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks cites this paper.

Understanding Space Is Rocket Science -- Only Top Reasoning Models Can Solve Spatial Understanding Tasks Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 63

Resolution
unresolved
no resolver link, observed 2026-08-05T11:56:10.921205Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T11:56:10.921205Z digest=sha256:6f641f2ca68bcd61b23e3882e8f50c7c015d0cd5e392f559cbc41c11975ed0ec

Observation dd397fc2-4ae1-42b8-ba90-65c061c5ff18 · inbound

Agentic Learner with Grow-and-Refine Multimodal Semantic Memory cites this paper.

Agentic Learner with Grow-and-Refine Multimodal Semantic Memory Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 30

Resolution
verified exact
arxiv_id, observed 2026-05-17T04:29:01.594708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-17T04:27:40.232015Z digest=sha256:6df5ca2f8e6637c018eedc5053f8e37c351f477a9dbb2e59721efbded8f96342

Observation 61285ced-81ef-45f4-aa01-93533de21893 · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 60

Resolution
verified exact
arxiv_id, observed 2026-05-10T16:09:05.382165Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:70f4ca4891f87eac52918acf19399e06712ba37a6ca33b298288ac076f082745

Observation d7fc9b8a-1cc8-4058-8cb7-618ac66d9d06 · inbound

Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension cites this paper.

Visual Para-Thinker: Divide-and-Conquer Reasoning for Visual Comprehension Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 20

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T03:30:33.014248Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-16T03:27:53.694506Z digest=sha256:4f4c850d80394fb24c48a8522a5ca8d2e82f716b0812fa6deb28bfc405879a86

Observation 8d35bc86-3a3c-4cfc-be9f-1b1a390b3d28 · inbound

Revisiting Model Stitching In the Foundation Model Era cites this paper.

Revisiting Model Stitching In the Foundation Model Era Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 45

Resolution
unresolved
no resolver link, observed 2026-07-14T22:19:38.973091Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-14T22:19:38.973091Z digest=sha256:8a975870f5bc83da4b7dce97ebf0c5c8541adc9fe4f54259a9e095a1fe3683ed

Observation b0d5d9e9-0238-4fb9-bc1b-1f2f903be8a2 · inbound

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models cites this paper.

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 83

Resolution
verified exact
arxiv_id, observed 2026-05-11T09:25:59.280081Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T16:03:15.222571Z digest=sha256:fe878e538d8eea4b05f1e57bb549f382fdab1dfd5bb11e2d81f496afbc37cdaf

Observation 88b072a1-0dd7-4ab6-9610-fcb7454daba0 · inbound

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models cites this paper.

Cognitive Pivot Points and Visual Anchoring: Unveiling and Rectifying Hallucinations in Multimodal Reasoning Models Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 95

Resolution
unresolved
no resolver link, observed 2026-07-12T22:48:45.647588Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:48:45.647588Z digest=sha256:6f8da8fa73a3481a2ccdc126382ecffd78e693dd578e77a3c92e6920d9c2e454

Observation 52b91b38-6d68-4b2f-a4a7-ca77ac00eb1f · inbound

Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation cites this paper.

Gaslight, Gatekeep, V1-V3: Early Visual Cortex Alignment Shields Vision-Language Models from Sycophantic Manipulation Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 45

Resolution
verified exact
arxiv_id, observed 2026-05-10T13:10:26.177537Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-10T13:09:35.407790Z digest=sha256:1ddd3388866658ffa2bdb18c4daddce044c6433f874cb7abaca5cad1a9483969

Observation 1cea5ba8-9582-406c-91f6-83972d727325 · inbound

SinkRouter: Sink-Aware Routing for Efficient Long-Context Decoding in Large Language and Multimodal Models cites this paper.

SinkRouter: Sink-Aware Routing for Efficient Long-Context Decoding in Large Language and Multimodal Models Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 26

Resolution
verified exact
arxiv_id, observed 2026-05-10T07:57:15.441339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-10T07:56:05.583390Z digest=sha256:dda21d3db60c76b959c3b3787d5c7067f34a3bc2706b783b825310c6eb6a9273

Observation b574e902-cc1f-4634-8f37-eccd69075fc3 · inbound

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm cites this paper.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-11T13:51:05.103821Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-09T23:56:02.856878Z digest=sha256:45a9a2dcc076980b29ca1d41d647519b0db7b2d5147eb57b153cedd47e97c0fb

Observation 5bf38efe-c0b3-4175-ae13-d0f1cb80751c · inbound

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm cites this paper.

The Expense of Seeing: Attaining Trustworthy Multimodal Reasoning Within the Monolithic Paradigm Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 14

Resolution
verified exact
arxiv_id, observed 2026-05-22T10:36:24.988618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-05-22T10:35:50.838150Z digest=sha256:dc7bd9f79928f425187930c24c5ea5c7b288ee9cf4d06fa274f73e503e7dd855

Observation 4ad5867c-504a-44f8-be22-bb2dda432282 · inbound

HypEHR: Hyperbolic Modeling of Electronic Health Records for Efficient Question Answering cites this paper.

HypEHR: Hyperbolic Modeling of Electronic Health Records for Efficient Question Answering Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 251

Resolution
verified exact
arxiv_id, observed 2026-05-09T23:54:45.492581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-05-09T23:51:47.724033Z digest=sha256:103e6ad20034548537c455d6df21338769d860f5ddbe70f7e8d79b7dde936486

Observation e12fcc5e-a825-45fb-be72-40741c6758aa · inbound

VisualNeedle: Benchmarking Active Visual Search in Information-Dense Scenes cites this paper.

VisualNeedle: Benchmarking Active Visual Search in Information-Dense Scenes Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 3

Resolution
metadata mismatch
arxiv_id, observed 2026-06-29T22:23:59.757075Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-29T22:14:00.535472Z digest=sha256:6d7f25525b510f509b4b80bdfc05d852db6eb66e7231aff722d2ea193b0829e2

Observation 589070f9-af55-4bd1-94c9-a01ecd1e9b50 · inbound

VESTA: Visual Exploration with Statistical Tool Agents cites this paper.

VESTA: Visual Exploration with Statistical Tool Agents Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 46

Resolution
verified exact
arxiv_id, observed 2026-07-01T19:56:10.468376Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-28T21:58:11.339217Z digest=sha256:3fb99b677145db8a334f5adbf0e24c37d4f9e0f1c5244e6f576584b1e977dd26

Observation 0d379c7f-0032-4471-934d-5751010ed5f5 · inbound

Fine-tuning Multi-modal LLMs with ART: Art-based Reinforcement Training cites this paper.

Fine-tuning Multi-modal LLMs with ART: Art-based Reinforcement Training Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-07-03T09:07:48.143928Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=pdf_text observed=2026-06-27T10:30:51.490047Z digest=sha256:e164013415d11000729e9b67033036c2ea6ba1cc5b98cef1f9380ead92261ed3

Observation 72f3f772-e9ff-400a-af7d-525651eb08f5 · inbound

Bridging the Usability Gap: Lessons from Interpreting Studies for Machine Interpreting Design cites this paper.

Bridging the Usability Gap: Lessons from Interpreting Studies for Machine Interpreting Design Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 87

Resolution
verified exact
arxiv_id, observed 2026-06-27T03:40:27.972918Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-27T03:34:37.373928Z digest=sha256:4b08f5195cae8efe1f2e669c475d1316922454980ded046acfd280a087f227a6

Observation 79206ccc-5d08-40bb-bc08-6dacc2b386d8 · inbound

Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning cites this paper.

Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 27

Resolution
metadata mismatch
arxiv_id, observed 2026-07-03T23:49:02.268563Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-09T06:31:02.800959+00:00.

source=arxiv_source observed=2026-06-26T21:47:51.437284Z digest=sha256:00f3520f8f23fc04920e6483b9863f5b84ee223c020db3b43b6ed8c87dadd65c

Observation 1c4639ee-9939-43cf-8533-2a3508e43a68 · inbound

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences cites this paper.

DELTAVID: Enhancing Fine-Grained Spatiotemporal Perception with Cross-Video Differences Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 48

Resolution
unresolved
no resolver link, observed 2026-07-12T11:31:14.532101Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T11:31:14.532101Z digest=sha256:350a09eb30474def985451023757b00c6f1ee318074b180ec5488db15cbd7629

Observation 9c1a1b0b-26a0-4323-8fdc-0a01a5581cfa · inbound

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation cites this paper.

VAD: Attributing Visual Evidence for Target Reconstruction in Multimodal On-Policy Distillation Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 88

Resolution
unresolved
no resolver link, observed 2026-07-31T02:56:58.277821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-31T02:56:58.277821Z digest=sha256:26c5acfd6156ce6e34aee0699024ed03d457ac6bffd5c9eb3333a7bac4aa1c85

Observation 8e51ee67-2932-4e6c-a1ff-4b3e0843eb3e · inbound

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping cites this paper.

The Low Frequency Trap: Video Language Models Fail at Simple Event Bookkeeping Eyes Wide Shut? Exploring the Visual Shortcomings of Multimodal LLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T04:24:54.770078Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T04:24:54.770078Z digest=sha256:4ccfdf071b0967f4a318b54ad3a6dc224740cadef37b39d8f8c75581ad3d789f