Pith. sign in

Paper Citation Record · LEDGER

On the robustness of multimodal language model towards distractions

As of 8 August 2026, this Paper Citation Record lists 37 of 37 outbound references and 2 inbound Pith citation observations for arXiv:2502.09818.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2502.09818 v1

Coverage vector

measured 37 of 37 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T20:25:58.596793Z

measured 39 of 39 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T01:00:27.572834Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-18T11:11:17.919527Z

Reference resolution

37 of 37 outbound references displayed

  • verified exact0
  • verified fuzzy23
  • unresolved13
  • parse uncertain1
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d37f96c1-4c41-4b8c-919a-0a402e321173 · outbound

This paper cites Vqa: Visual question answering.

On the robustness of multimodal language model towards distractions Vqa: Visual question answering

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.350990Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.384098Z digest=sha256:1527a822aea3a7e00d1154930aec8501be628a9a38364f75e55a861dbba6c35a

Observation c5f01393-06f9-4c4a-8a5f-d819ccc290c6 · outbound

This paper cites Choquette- Choo, Matthew Jagielski, Irena Gao, Anas Awadalla, Pang Wei Koh, Daphne Ippolito, Katherine Lee, Florian Tramer, and Ludwig Schmidt.

On the robustness of multimodal language model towards distractions Choquette- Choo, Matthew Jagielski, Irena Gao, Anas Awadalla, Pang Wei Koh, Daphne Ippolito, Katherine Lee, Florian Tramer, and Ludwig Schmidt

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.331941Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.391159Z digest=sha256:d7a427c4d418d90bd1bb63cf684a955de8111e286432339714c263cb501fe5d7

Observation b75241f2-2493-4802-9562-c21873821f9d · outbound

This paper cites How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites.

On the robustness of multimodal language model towards distractions How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.396754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.396754Z digest=sha256:dd76bd52da0998f72c60f77213cbe64ad1d6ad216ddcc634ca0001ff91a85fc9

Observation dffe2634-4015-4213-af49-7a35f2e7c4c4 · outbound

This paper cites HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding.

On the robustness of multimodal language model towards distractions HALC: Object Hallucination Reduction via Adaptive Focal-Contrast Decoding

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.403744Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.403744Z digest=sha256:e41ff77864faf15dd2577f0c14b01fdbc92b7f55585326c1a4b6c39cc43179f7

Observation 1e68e2c4-f4ec-469a-9f33-44081099c8c6 · outbound

This paper cites Instructblip: Towards general- purpose vision-language models with instruction tuning,.

On the robustness of multimodal language model towards distractions Instructblip: Towards general- purpose vision-language models with instruction tuning,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.410363Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.410363Z digest=sha256:6c2d21551e8bebac932c2cc5b32b5523e21e68f7549672af327c982074809a9b

Observation 0cede9a2-b671-4098-82ae-3f68818a795a · outbound

This paper cites How Robust is Google's Bard to Adversarial Image Attacks?.

On the robustness of multimodal language model towards distractions How Robust is Google's Bard to Adversarial Image Attacks?

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.416124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.416124Z digest=sha256:5fa913a480c4eb9055b6a715612d9bbaa633078154e4c067d2ff80185fd7ae17

Observation 0fb5e3ff-151b-4fce-97c8-024e0321cc11 · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

On the robustness of multimodal language model towards distractions MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.422649Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.422649Z digest=sha256:44bd898c7b17aea2ed1c4cf4f78208ecb375aaba844bbd63a904b8bc0b49aced

Observation 32aa141f-b92e-476b-97fd-c2eb6c6fdb1d · outbound

This paper cites Phi3v-finetuning, 2023.

On the robustness of multimodal language model towards distractions Phi3v-finetuning, 2023

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.300025Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.429141Z digest=sha256:725687bd2f1c4a15700c346cebba2eec6b075722d954a61fbe44a17c3ad25248

Observation 19b751bb-1b39-44f4-8568-0f24ab94da08 · outbound

This paper cites HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models.

On the robustness of multimodal language model towards distractions HallusionBench: An Advanced Diagnostic Suite for Entangled Language Hallucination and Visual Illusion in Large Vision-Language Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.434448Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.434448Z digest=sha256:c34d75abbac8022d88a781746cc0c6a2472504b1790016e71df1a21d8477227b

Observation a8e5ac55-2ec6-4eed-b9b3-ba4cff84de8e · outbound

This paper cites Cogvlm2: Visual language models for image and video understanding, 2024.

On the robustness of multimodal language model towards distractions Cogvlm2: Visual language models for image and video understanding, 2024

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.281755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.440591Z digest=sha256:dd407fb8954ca9b95c29645633ddc66cdbd8b9e456c48e866fdca103769c1b26

Observation 16072865-e658-44f3-ad2a-1c9ca4557ae9 · outbound

This paper cites Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages.

On the robustness of multimodal language model towards distractions Large Multilingual Models Pivot Zero-Shot Multimodal Learning across Languages

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.446123Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.446123Z digest=sha256:651b3e7f0f98e31e8a4f6b4655b193e8507f8cb9f9c54ab15cb7b0bd0db294ca

Observation bbc59deb-19da-45ee-865c-468d971daa7f · outbound

This paper cites Gqa: A new dataset for real-world visual reasoning and compositional question answering.

On the robustness of multimodal language model towards distractions Gqa: A new dataset for real-world visual reasoning and compositional question answering

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.262535Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.451812Z digest=sha256:536bceaacd60b11d5d7ec60a281590ee32d6fbec11cce3c679a96613d698ad2f

Observation 4a7358af-0757-499b-bb12-dfbab7ebdd5b · outbound

This paper cites Open- clip, 2021.

On the robustness of multimodal language model towards distractions Open- clip, 2021

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.458116Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.458116Z digest=sha256:420a3447cd729d019dc4107b9adb68c85c61be2b0140fa0675b4bcee3051f787

Observation a3ac1863-e170-4d35-8487-75b3139b806d · outbound

This paper cites Adversarial examples for evaluating math word problem solvers.

On the robustness of multimodal language model towards distractions Adversarial examples for evaluating math word problem solvers

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.230192Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.464365Z digest=sha256:c7b12668e8510c6f6353971c5e53fe73078655a8c14da5a6538e69d4a0a0c4b2

Observation 69848a13-51bb-4a10-a771-a3e89c5162d9 · outbound

This paper cites LISA: Reasoning Segmentation via Large Language Model.

On the robustness of multimodal language model towards distractions LISA: Reasoning Segmentation via Large Language Model

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.469758Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.469758Z digest=sha256:fba99858d2207224af7b84902803bfe72917fc0b7c5c9f69cd2afb6337d74596

Observation b327ad1f-67ed-4043-8270-8c0e367b89bb · outbound

This paper cites Evaluating object hallucination in large vision-language models.

On the robustness of multimodal language model towards distractions Evaluating object hallucination in large vision-language models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.476188Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.476188Z digest=sha256:279d88e23eee5b8ce49b7d8ee1bd22271a933a71a3b25642fa19039bd9cc9a14

Observation 010bf48e-4ce5-44fe-a456-0a2d80a7dd25 · outbound

This paper cites Visual instruction tuning.

On the robustness of multimodal language model towards distractions Visual instruction tuning

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.200014Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.481824Z digest=sha256:b3c8780be5c9337058784e4e8684347541b6a2a730afb5d373c8ad3140e2613f

Observation 2394962e-aa9d-4613-86c3-49bfbe126f3e · outbound

This paper cites MMBench: Is Your Multi-modal Model an All-around Player?.

On the robustness of multimodal language model towards distractions MMBench: Is Your Multi-modal Model an All-around Player?

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.486806Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.486806Z digest=sha256:0c6bbd5e15793601e2d464d45c58e10cb45396290321079e81094e6256c9fb6b

Observation 38ffbaa5-1e8d-4562-aeaa-3b5fac7cd4cd · outbound

This paper cites Learn to explain: Multimodal reasoning via thought chains for science question answering.

On the robustness of multimodal language model towards distractions Learn to explain: Multimodal reasoning via thought chains for science question answering

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.177301Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.493078Z digest=sha256:35e8308c71a42333e43423a5daab7b4a3f461019c05cd5125025786d07ccf9f9

Observation cd7b49c2-a6e2-43bb-9766-1022caeef332 · outbound

This paper cites Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts, 2024.

On the robustness of multimodal language model towards distractions Mathvista: Evaluating mathemat- ical reasoning of foundation models in visual contexts, 2024

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.157478Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.498077Z digest=sha256:7596a635ce2cfa4740c56636dd619f97feda7b7a08cb9c5c4dd3a806e37347b6

Observation 06812f5e-2460-458d-828b-3204904a853d · outbound

This paper cites Understanding zero-shot adversarial robust- ness for large-scale models.

On the robustness of multimodal language model towards distractions Understanding zero-shot adversarial robust- ness for large-scale models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.136750Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.503335Z digest=sha256:63f2281f3cd0ae6d9a72153c6a999c2ab0ca5003066231c4280736bebfaae0c5

Observation 3b5c2d96-6aef-4ffc-a56e-fd306bc2d6c2 · outbound

This paper cites Gpt-4v(ision) system card.

On the robustness of multimodal language model towards distractions Gpt-4v(ision) system card

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.508674Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.508674Z digest=sha256:69c5eaf69f9d0bb1ad22ca7f237e5f62ba945471cef22c5067f02e40a8c96f4c

Observation cc9d91b8-b302-4cf8-aba5-62fb66f13ecf · outbound

This paper cites Gpt-3.5 turbo.

On the robustness of multimodal language model towards distractions Gpt-3.5 turbo

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.102629Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.514303Z digest=sha256:1c9a755a44361392815289d3bec989d15f83bfa095a1d4e072682533588f85a4

Observation 403be801-f927-4acd-b089-efe41a4ba526 · outbound

This paper cites Are nlp models really able to solve simple math word problems? In NAACL-HLT, 2021.

On the robustness of multimodal language model towards distractions Are nlp models really able to solve simple math word problems? In NAACL-HLT, 2021

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.081273Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.519582Z digest=sha256:b2aca9d2f4b0a26364174b755a32c147287f3b6afeaf33a938ca15233bc1db00

Observation b32d299b-6207-43d8-96f0-8fead312d049 · outbound

This paper cites Homepage.

On the robustness of multimodal language model towards distractions Homepage

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.061065Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.525855Z digest=sha256:1a9899ffc33d4a0e66c97ffa29a0319639559152b3873a1e0e2b889774b20387

Observation 3439da96-30d7-4e0b-b9f1-8b278ed8d560 · outbound

This paper cites Visual adversarial examples jailbreak aligned large language models.

On the robustness of multimodal language model towards distractions Visual adversarial examples jailbreak aligned large language models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:59.008222Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.539490Z digest=sha256:411918a5bbd6d50b9e3b461d025e809579afc687c23608188728c911a8e14204

Observation f6299b28-b9e6-4049-91a2-2ea929020757 · outbound

This paper cites Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer.

On the robustness of multimodal language model towards distractions Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:58.986614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.545301Z digest=sha256:e1592ef029470c463bd569ee5e38a8e20abdc9140ef69df67db22995cabfa7b4

Observation a5ce974a-f315-43a0-a3e6-d0e34f3137da · outbound

This paper cites On the adversarial robustness of multi-modal foundation models, 2023.

On the robustness of multimodal language model towards distractions On the adversarial robustness of multi-modal foundation models, 2023

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:58.967524Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.550995Z digest=sha256:4bcabafb1f2d2daa1081960548c03122bfa634e8a78837fb00e2a0eb8332eda6

Observation f81d975d-aed2-4629-8323-c292deccde64 · outbound

This paper cites Robust clip: Unsupervised ad- versarial fine-tuning of vision embeddings for robust large vision-language models.

On the robustness of multimodal language model towards distractions Robust clip: Unsupervised ad- versarial fine-tuning of vision embeddings for robust large vision-language models

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:58.948610Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.556555Z digest=sha256:a6f3f5997292bf8678c412255ac5e23b277ce189bd4b5a46990242f4012c42fe

Observation 87b13a6c-1f33-46f3-b4c0-dc1caf2df972 · outbound

This paper cites Large language models can be easily distracted by irrelevant context, 2023.

On the robustness of multimodal language model towards distractions Large language models can be easily distracted by irrelevant context, 2023

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:58.928046Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.562026Z digest=sha256:f32cf1291c76f90287ce06db94e1f681165cb213133b0f8f9ce61f3b8023f66f

Observation 4a016203-6941-4686-b224-43f566834c56 · outbound

This paper cites Towards vqa models that can read.

On the robustness of multimodal language model towards distractions Towards vqa models that can read

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:58.908306Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.568109Z digest=sha256:ac2bf774a1e931e4bd7f93513cd632c85207aa615246a4dacfe6f54b78ad3106

Observation a31188e5-e4eb-4d69-bb43-306c9759323f · outbound

This paper cites Unsplash api.

On the robustness of multimodal language model towards distractions Unsplash api

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:58.889018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.573558Z digest=sha256:f6ddae432e3bcdd8564c23359f72b50e473b1cf56a87027c96f546dada614c47

Observation 88d427a4-54cc-4ea6-92a4-42de4e6d29f4 · outbound

This paper cites Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024.

On the robustness of multimodal language model towards distractions Qwen2-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:58.870325Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.580371Z digest=sha256:9cc6f3179d5c05ce0e43c1bd35f562e94c30e91ed9ee4dd47ab79f46409679f3

Observation 511b0599-cc63-4731-94ce-6bdcecf8d2ce · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

On the robustness of multimodal language model towards distractions MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T20:25:58.585610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-07T20:25:58.585610Z digest=sha256:bd459880398e9faab87ad39923fcdfe77eb86ffde96d7afa8c7144dfb288da30

Observation c0365983-3f1b-44a0-9ff2-b663f50f370d · outbound

This paper cites On evaluating ad- versarial robustness of large vision-language models, 2023.

On the robustness of multimodal language model towards distractions On evaluating ad- versarial robustness of large vision-language models, 2023

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:58.851745Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.591386Z digest=sha256:5b24d91f17682b9bda273930d8cd08dbfbcc364525abfde85e66c2f97a611f53

Observation 02225cbf-15c6-4914-b4b6-8fdb52c6be58 · outbound

This paper cites Analyzing and mitigating object hallucination in large vision-language models.

On the robustness of multimodal language model towards distractions Analyzing and mitigating object hallucination in large vision-language models

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T20:25:58.827053Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.596793Z digest=sha256:d3d8e963735012481d9537c2cd43513742483eaa101ca106203c75f586cf1fd1

Observation f4d9cd83-44e4-40e9-bdf3-23e12f2c026c · outbound

This paper cites an unresolved cited work.

On the robustness of multimodal language model towards distractions Unresolved cited work

Reference 2024

Resolution
parse uncertain
raw_fallback, observed 2026-08-07T20:25:59.041438Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-08-07T20:25:58.533546Z digest=sha256:f35cfe62fc5e58d02d6acec4c1f56e881b9f47f8c50ef40e1b435e7dcb78d49a

Pith citing papers

Observation 1d1532fb-9e7f-4bea-bd60-4c445ac19072 · inbound

Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability cites this paper.

Can Large Multimodal Models Actively Recognize Faulty Inputs? A Systematic Evaluation Framework of Their Input Scrutiny Ability On the robustness of multimodal language model towards distractions

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-06T01:00:27.572834Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T01:00:27.572834Z digest=sha256:3db6559060ee8cd8fc9e1fba8e9d4771dd05f8cd15ad967583eab1f5d8682d26

Observation a79769f4-ea2d-41a4-b959-8c00f88e3b89 · inbound

When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models cites this paper.

When Silence Matters: The Impact of Irrelevant Audio on Text Reasoning in Large Audio-Language Models On the robustness of multimodal language model towards distractions

Reference 29

Resolution
verified exact
arxiv_id, observed 2026-05-18T11:11:17.922848Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-18T11:08:08.916893Z digest=sha256:2da6b337a24dc3447632d8af13c4c7bd4783a44c02f7fd37052f5a9c1c19e3cf