Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T20:25:58.596793Z
Paper Citation Record · LEDGER
As of 8 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 2 inbound Pith citation observations for arXiv:2502.09818.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T20:25:58.596793Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-06T01:00:27.572834Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-18T11:11:17.919527Z
37 of 37 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d37f96c1-4c41-4b8c-919a-0a402e321173 · outbound
On the robustness of multimodal language model towards distractions Vqa: Visual question answering
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation c5f01393-06f9-4c4a-8a5f-d819ccc290c6 · outbound
On the robustness of multimodal language model towards distractions Choquette- Choo, Matthew Jagielski, Irena Gao, Anas Awadalla, Pang Wei Koh, Daphne Ippolito, Katherine Lee, Florian Tramer, and Ludwig Schmidt
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b75241f2-2493-4802-9562-c21873821f9d · outbound
On the robustness of multimodal language model towards distractions How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dffe2634-4015-4213-af49-7a35f2e7c4c4 · outbound
On the robustness of multimodal language model towards distractions HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e68e2c4-f4ec-469a-9f33-44081099c8c6 · outbound
On the robustness of multimodal language model towards distractions Instructblip: Towards general- purpose vision-language models with instruction tuning,
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0cede9a2-b671-4098-82ae-3f68818a795a · outbound
On the robustness of multimodal language model towards distractions How Robust is Google's Bard to Adversarial Image Attacks?
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0fb5e3ff-151b-4fce-97c8-024e0321cc11 · outbound
On the robustness of multimodal language model towards distractions MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 32aa141f-b92e-476b-97fd-c2eb6c6fdb1d · outbound
On the robustness of multimodal language model towards distractions Phi3v-finetuning, 2023
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 19b751bb-1b39-44f4-8568-0f24ab94da08 · outbound
On the robustness of multimodal language model towards distractions HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a8e5ac55-2ec6-4eed-b9b3-ba4cff84de8e · outbound
On the robustness of multimodal language model towards distractions Cogvlm2: Visual language models for image and video understanding, 2024
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 16072865-e658-44f3-ad2a-1c9ca4557ae9 · outbound
On the robustness of multimodal language model towards distractions Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bbc59deb-19da-45ee-865c-468d971daa7f · outbound
On the robustness of multimodal language model towards distractions Gqa: A new dataset for real-world visual reasoning and compositional question answering
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4a7358af-0757-499b-bb12-dfbab7ebdd5b · outbound
On the robustness of multimodal language model towards distractions Open- clip, 2021
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a3ac1863-e170-4d35-8487-75b3139b806d · outbound
On the robustness of multimodal language model towards distractions Adversarial examples for evaluating math word problem solvers
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 69848a13-51bb-4a10-a771-a3e89c5162d9 · outbound
On the robustness of multimodal language model towards distractions LISA: Reasoning Segmentation via Large Language Model
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b327ad1f-67ed-4043-8270-8c0e367b89bb · outbound
On the robustness of multimodal language model towards distractions Evaluating object hallucination in large vision-language models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 010bf48e-4ce5-44fe-a456-0a2d80a7dd25 · outbound
On the robustness of multimodal language model towards distractions Visual instruction tuning
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 2394962e-aa9d-4613-86c3-49bfbe126f3e · outbound
On the robustness of multimodal language model towards distractions MMBench: Is Your Multi-modal Model an All-around Player?
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 38ffbaa5-1e8d-4562-aeaa-3b5fac7cd4cd · outbound
On the robustness of multimodal language model towards distractions Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation cd7b49c2-a6e2-43bb-9766-1022caeef332 · outbound
On the robustness of multimodal language model towards distractions Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts, 2024
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 06812f5e-2460-458d-828b-3204904a853d · outbound
On the robustness of multimodal language model towards distractions Understanding zero-shot adversarial robust- ness for large-scale models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3b5c2d96-6aef-4ffc-a56e-fd306bc2d6c2 · outbound
On the robustness of multimodal language model towards distractions Gpt-4v(ision) system card
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cc9d91b8-b302-4cf8-aba5-62fb66f13ecf · outbound
On the robustness of multimodal language model towards distractions Gpt-3.5 turbo
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 403be801-f927-4acd-b089-efe41a4ba526 · outbound
On the robustness of multimodal language model towards distractions Are nlp models really able to solve simple math word problems? In NAACL-HLT, 2021
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation b32d299b-6207-43d8-96f0-8fead312d049 · outbound
On the robustness of multimodal language model towards distractions Homepage
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 3439da96-30d7-4e0b-b9f1-8b278ed8d560 · outbound
On the robustness of multimodal language model towards distractions Visual adversarial examples jailbreak aligned large language models
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f6299b28-b9e6-4049-91a2-2ea929020757 · outbound
On the robustness of multimodal language model towards distractions Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a5ce974a-f315-43a0-a3e6-d0e34f3137da · outbound
On the robustness of multimodal language model towards distractions On the adversarial robustness of multi-modal foundation models, 2023
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f81d975d-aed2-4629-8323-c292deccde64 · outbound
On the robustness of multimodal language model towards distractions Robust clip: Unsupervised ad- versarial fine-tuning of vision embeddings for robust large vision-language models
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 87b13a6c-1f33-46f3-b4c0-dc1caf2df972 · outbound
On the robustness of multimodal language model towards distractions Large language models can be easily distracted by irrelevant context, 2023
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 4a016203-6941-4686-b224-43f566834c56 · outbound
On the robustness of multimodal language model towards distractions Towards vqa models that can read
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation a31188e5-e4eb-4d69-bb43-306c9759323f · outbound
On the robustness of multimodal language model towards distractions Unsplash api
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 88d427a4-54cc-4ea6-92a4-42de4e6d29f4 · outbound
On the robustness of multimodal language model towards distractions Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 511b0599-cc63-4731-94ce-6bdcecf8d2ce · outbound
On the robustness of multimodal language model towards distractions MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c0365983-3f1b-44a0-9ff2-b663f50f370d · outbound
On the robustness of multimodal language model towards distractions On evaluating ad- versarial robustness of large vision-language models, 2023
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 02225cbf-15c6-4914-b4b6-8fdb52c6be58 · outbound
On the robustness of multimodal language model towards distractions Analyzing and mitigating object hallucination in large vision-language models
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation f4d9cd83-44e4-40e9-bdf3-23e12f2c026c · outbound
On the robustness of multimodal language model towards distractions Unresolved cited work
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.
Observation 1d1532fb-9e7f-4bea-bd60-4c445ac19072 · inbound
Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability On the robustness of multimodal language model towards distractions
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a79769f4-ea2d-41a4-b959-8c00f88e3b89 · inbound
When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models On the robustness of multimodal language model towards distractions
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.