Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:08:51.495834Z
Paper Citation Record · LEDGER
As of 22 August 2026, this Paper Citation Record lists 54 of 54 outbound references and 8 inbound Pith citation observations for arXiv:2411.17491.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-12T12:08:51.495834Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-15T15:05:19.272138Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-07-08T05:54:33.575217Z
54 of 54 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation de162673-5c57-4daa-9aa8-86f1473d7bd1 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models GPT-4 Technical Report
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5b22506-24a0-4e0f-a8ac-5e58e5b63171 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Flamingo: a visual language model for few-shot learning
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 087802d2-6fab-4b70-b3da-6d7ddc6b265f · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Hallucination of Multimodal Large Language Models: A Survey
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c5281822-b368-42a7-a429-9a55a42a0703 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Understanding Information Storage and Transfer in Multi-modal Large Language Models
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation beb9b1f7-2603-43fe-9f23-28c01888eb24 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 892eb2f5-f6e5-4c20-85a2-2f1fd3b6fbf7 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2c5e41ca-7797-44e9-9092-f701b65660a5 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models InternVL: Scaling up vision foundation models and aligning for generic visual-linguistic tasks
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation af800eb7-bc7a-49ae-ad87-0c72d78b3f4c · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models What Does BERT Look At? An Analysis of BERT's Attention
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1c8bde47-d19b-45bd-9b00-ac98c8ef9461 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Towards automated circuit discovery for mechanistic interpretability
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d6ffd854-c6f7-4f43-8c74-f458c8a0410a · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models The Llama 3 Herd of Models
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 11c432f5-df61-4d2d-8652-c4014c85c10b · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 762f1bae-143f-4988-b852-fe6dfba05150 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Transformer Feed-Forward Layers Are Key-Value Memories
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 69994d22-d342-4955-87e9-fc3bb30f59b0 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Dissecting Recall of Factual Associations in Auto-Regressive Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e7439443-0c2a-4d36-8655-d036e87cd50b · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Chat- gpt outperforms crowd workers for text-annotation tasks
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a6fe1baa-c504-4b3e-b6e3-a84e9024d6cb · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models The Unreasonable Ineffectiveness of the Deeper Layers
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 36b2e311-9696-4ac9-b50c-540d6696c5be · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Conceptgraphs: Open-vocabulary 3d scene graphs for per- ception and planning
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5ee653b1-5446-478e-86a5-0f086855a170 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Is chatgpt better than human annotators? potential and limitations of chatgpt in explaining implicit hate speech
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3d0eefc8-f3f7-4c04-9799-4d0f9b678951 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Mistral 7B
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 23d4f89a-e093-4483-952c-f7d56133c2f8 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Interpreting and Editing Vision-Language Representations to Mitigate Hallucinations
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f30c3830-d93c-40f3-9e8d-a8f0597fa5d2 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Segment any- thing
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 0d8990f3-234b-4afc-95c1-294ea1ca8b52 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1b0f2264-ffa6-4b98-8eba-dfb359cab21d · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models LLM-grounded Video Diffusion Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4d141e33-a30a-4f55-a70a-1c097ee82b25 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Microsoft coco: Common objects in context
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9accea1c-b706-46d7-8a4e-b15e7e7dcba2 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Improved baselines with visual instruction tuning
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation fe2364bd-9b3a-4fa7-a5c8-5f8c273d67e4 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Visual instruction tuning
Reference 25
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 134b4812-cce6-421a-be8a-8f15f0430de6 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Clip-driven universal model for organ segmentation and tumor detection
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd0fc225-4d62-4fd3-a26b-b4589b72f26b · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models OK-Robot: What Really Matters in Integrating Open-Knowledge Models for Robotics
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d1d4dd65-50d5-4040-a114-327c57766e75 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Locating and editing factual associations in gpt
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 73d73144-6432-4148-9f8a-c10b50ba680b · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Towards Interpreting Visual Information Processing in Vision-Language Models
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5281fb7e-973f-4dee-b8bf-c06ac3d31b2f · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Interpreting gpt: The logit lens
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b41ac9f2-1cee-4e82-b4c6-0217cac9eb25 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Learning transferable visual models from natural language supervi- sion
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d516d450-f816-4a82-8e34-11933065aa28 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models A multimodal automated interpretability agent
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation d04c8566-f351-47ef-a2c9-04baecb01423 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models ChartMimic: Evaluating LMM's Cross-Modal Reasoning Capability via Chart-to-Code Generation
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 73eb3bb7-fda1-4246-8e3a-b3aab52b2418 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models DriveVLM: The Convergence of Autonomous Driving and Large Vision-Language Models
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4f574b5b-25e9-4d99-a0f9-906b87061d33 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bda0570c-9744-4e3d-b21d-977a0f286b40 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 36
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 51702afa-4fef-46f5-88c0-487e4462ae75 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Chain-of-thought prompting elicits reasoning in large lan- guage models
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 0281bca4-c325-4755-a8ac-af03796a52e1 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Qwen2 Technical Report
Reference 38
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8a863dfb-10ab-4558-80f1-393dde893a8e · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models EVF-SAM: Early Vision-Language Fusion for Text-Prompted Segment Anything Model
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bfb880ad-6c13-4597-8f83-50af331a3d1a · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Yes” and “No
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f51b07f5-21bf-4cd0-808f-b3d1e1c630c7 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Detect only tangible objects that can be interacted with
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9f2c0342-4454-4093-86ec-12f7d0279d76 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Unresolved cited work
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 5da0c476-abd1-4512-8179-5f2a724f5722 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models If half of the *physical objects* in the predicted caption are also in the groundtruth caption, the precision would be 0.5
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 9d2d3c9b-50c3-4cf0-8050-b26dd699507d · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models fine detail
Reference 44
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation a82c2b7e-f90a-4250-b4a5-8f04be1fbf0c · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Unresolved cited work
Reference 45
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 25761686-1e43-4a08-a42a-2be8d0e0fcc1 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Unresolved cited work
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 3653dfb8-23e9-403a-b553-f87fb5ebd7c6 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Unresolved cited work
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 73d9ce69-4a03-4532-802f-d308c8b4c901 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Unresolved cited work
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation ac485285-edda-4d31-bf3c-fc2d54092387 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Unresolved cited work
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 78f7737a-3741-45b4-a557-3f4397f5e13a · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Unresolved cited work
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 2c1ee46b-1ae9-46a4-a143-a916844f4bc0 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Unresolved cited work
Reference 51
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 458bdb31-935d-4f0f-94e9-e4bdf654f928 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Unresolved cited work
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 24a70173-6466-4293-901b-0e2ab8cb358a · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Unresolved cited work
Reference 53
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation 941f93b4-b6a6-430f-820b-847cf1e2ffa9 · outbound
What's in the Image? A Deep-Dive into the Vision of Vision Language Models Please answer yes or no
Reference 54
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation b7d78133-b0b0-44da-9a83-3dc8da27fa44 · inbound
Investigating Mechanisms for In-Context Vision Language Binding What's in the Image? A Deep-Dive into the Vision of Vision Language Models
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fd8974b9-4d49-436a-8611-a22b476e4d89 · inbound
PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning What's in the Image? A Deep-Dive into the Vision of Vision Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42272f09-424f-4567-976e-8fde399d364d · inbound
Scale Your Instructions: Enhance the Instruction-Following Fidelity of Unified Image Generation Model by Self-Adaptive Attention Scaling What's in the Image? A Deep-Dive into the Vision of Vision Language Models
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fdb07771-8a63-4754-bea1-7ffeaf987dc0 · inbound
Counting to Four is still a Chore for VLMs What's in the Image? A Deep-Dive into the Vision of Vision Language Models
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation c3bcf70c-4763-4c29-890f-db5f2eb90ecc · inbound
Vision-Default, Prior-Override: Causal Mechanisms of Perception-Knowledge Conflict in Vision-Language Models What's in the Image? A Deep-Dive into the Vision of Vision Language Models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f099b476-2c82-4664-bae0-8de9a7610a42 · inbound
Analysis-by-Proxy: Localization Signals in VLMs Operating as Condition Encoders What's in the Image? A Deep-Dive into the Vision of Vision Language Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.
Observation f90f76b3-eed1-4a3b-a13a-947ff61d0e32 · inbound
In-Context Collapse in Vision-Language Models and How to Mitigate it? What's in the Image? A Deep-Dive into the Vision of Vision Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 72836731-785e-4f17-8230-0bacb9b81695 · inbound
Routing Is Least Learnable Where It Is Most Valuable: Bounds on Representation Routing for Web Agents What's in the Image? A Deep-Dive into the Vision of Vision Language Models
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.