Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:13:38.669736Z
Paper Citation Record · LEDGER
As of 17 August 2026, this Paper Citation Record lists 50 of 50 outbound references and 1 inbound Pith citation observation for arXiv:2505.10541.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:13:38.669736Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-17T20:53:15.920563Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-17T20:55:15.290281Z
50 of 50 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 8e65b50f-6a00-4907-aeca-e697b5d9b09d · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis write newline
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9e63ff9a-cf2b-4e2f-91ef-9f8f5f1bdbf2 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Claude 3 haiku: our fastest model yet
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c0766cfa-f7b5-4a79-b752-4cc7ea13fd01 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Y., Bhiwandiwalla, A., Tseng, S.-Y., Olson, M
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3d037445-bc1f-4ad0-95b6-eb6646212a53 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis F., G \'o mez, L., and Karatzas, D
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dcfe01a4-b34e-4c45-8b1f-319cbbd9dea6 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis H., Vora, S., Liong, V
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 439fd75e-fe80-46f7-8b96-def93d14fa28 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Are We on the Right Way for Evaluating Large Vision-Language Models?
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e6d834d3-9ba7-4424-8e7b-6145d80a4a3a · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Dress: Instructing large vision-language models to align and interact with humans via natural language feedback
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 95434e8c-a688-45e2-ba2d-08060f68889f · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 037abd1d-c667-4fd2-9824-58e20bdac74c · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis H., Yu, F., Wan, X., and Wang, B
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation b6de1c68-313a-4035-a964-4ac65c790d2c · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis A Challenger to GPT-4V? Early Explorations of Gemini in Visual Expertise
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13945699-3fac-49b0-8aaa-ca25d471b533 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis BLINK: Multimodal Large Language Models Can See but Not Perceive
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ddefc0fa-4660-4846-9098-4672dd91a1b8 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis ChatGLM: A Family of Large Language Models from GLM-130B to GLM-4 All Tools
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 623384f0-0c02-49c4-889a-0dac13b88e5e · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Hallusionbench: an advanced diagnostic suite for entangled language hallucination and visual illusion in large vision-language models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8059e091-9fd4-45f9-a8cb-c97e28baca0f · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis CMMU: A Benchmark for Chinese Multi-modal Multi-type Question Understanding and Reasoning
Reference 14
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e05b8189-4235-45af-b516-9166d7fce7c3 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis CogVLM2: Visual Language Models for Image and Video Understanding
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4162649e-6b6a-4df6-a251-80f50038dad4 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Self-Introspective Decoding: Alleviating Hallucinations for Large Vision-Language Models
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4418e467-ff88-4fec-9f4a-7231c02a0469 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis MANTIS: Interleaved Multi-Image Instruction Tuning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b42bf34f-0c98-41eb-9012-0ec11b9c3f63 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Are you smarter than a sixth grader? textbook question answering for multimodal machine comprehension
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e5144ad0-7b10-472d-8470-5b5cb7e2a691 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Obelics: An open web-scale filtered dataset of interleaved image-text documents
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f4188452-f5df-4d9d-939a-372a2929b963 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis The Curse of Multi-Modalities: Evaluating Hallucinations of Large Multimodal Models across Language, Visual, and Audio
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 3f50f6ce-c668-4786-bf7a-dfe77276435a · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5415a960-083b-48b7-8ea6-a74e4d64db68 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e174b456-dba6-4805-9865-ac6a25267585 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis VisualBERT: A Simple and Performant Baseline for Vision and Language
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 15a68c54-70e4-428d-beab-f690768d86f4 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Evaluating object hallucination in large vision-language models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d3d3625-3f5b-4bdf-84d0-70b160d07d33 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis X., Tian, P., Yin, C
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4e7cd78a-e8f8-4978-b5c4-17ebfe5758b7 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a838b0bf-022e-4f1e-a280-44925a1c28ed · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis MMDU: A Multi-Turn Multi-Image Dialog Understanding Benchmark and Instruction-Tuning Dataset for LVLMs
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56f88c9f-0cc8-4441-92be-cc0b506b312b · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Learn to explain: Multimodal reasoning via thought chains for science question answering
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 48c02c35-1732-4f56-938d-96686c26fb47 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Docvqa: A dataset for vqa on document images
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ec068e0e-6881-40ee-a553-be31aee913a8 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis MMIU: Multimodal Multi-image Understanding for Evaluating Large Vision-Language Models
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 39528761-a855-4a0a-b23f-2cfde322fefa · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Gpt-4o system card, 2024
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 64bdf4d0-5e00-49c9-a0c5-9e3fbce4c7da · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Perception test: A diagnostic benchmark for multimodal video models
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9e4ba0fa-2975-4356-a414-f8cd75e51486 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis U., Hezel, N., and Jung, K
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 814fce51-8e6c-4722-8abc-abe0fbc87f9d · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis EVA-CLIP: Improved Training Techniques for CLIP at Scale
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 778cad13-ab66-4424-84f3-0101bbc5fa4e · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Aligning large multimodal models with factually augmented RLHF
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1cd8ac92-a3bd-4306-9da4-355dff4b3e83 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Slidevqa: A dataset for document visual question answering on multiple images
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e3284111-3a0a-4561-b56b-95fd3705422a · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 09fd7815-4961-4df3-a3dd-66e6ed87c57e · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Attention is all you need
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 740686f6-9c71-45a0-bb11-87c78d0b8654 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Analyzing the Structure of Attention in a Transformer Language Model
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6808bb5d-b3d9-4bc9-a85d-98138846febf · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c105421a-acd4-4a3d-a25b-ea67ef4f8f3f · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69ca77c9-7cfb-49ff-b405-fd6937c4d8f9 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Needle in a multimodal haystack
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 66a3c391-a451-4848-b51d-fb017986a3a7 · outbound
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0fbb95f1-2556-4a89-a1a0-5148abc81d3f · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis MMIE: Massive Multimodal Interleaved Comprehension Benchmark for Large Vision-Language Models
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 66753a6f-05d7-4180-9997-5c9abc452796 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Qwen2 Technical Report
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 64d88026-b91b-4d5e-947d-0caff1f3308d · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V
Reference 46
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ef60c97a-fed6-4657-8c78-8620ef2887a4 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions
Reference 47
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6574eee5-562a-4fbb-b353-15b4f62a7736 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Mmmu: A massive multi-discipline multimodal understanding and reasoning benchmark for expert agi
Reference 48
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0aaaecd9-e590-4823-902f-79a487135dfb · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis Sigmoid loss for language image pre-training
Reference 49
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b1ab4e11-b315-46b1-9795-1d51107345e7 · outbound
Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis InternLM-XComposer-2.5: A Versatile Large Vision Language Model Supporting Long-Contextual Input and Output
Reference 50
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5cfd14b8-5ae3-4165-9053-3f6bcb53be59 · inbound
Attention Grounded Enhancement for Visual Document Retrieval Exploring Implicit Visual Misunderstandings in Multimodal Large Language Models through Attention Analysis
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.