Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:06:49.290784Z
Paper Citation Record · LEDGER
As of 16 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2505.05626.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T23:06:49.290784Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
26 of 26 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 3dace842-e939-43af-b761-3563d2ff32e4 · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Unresolved cited work
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 0d8090c1-07b4-4f97-ad1d-4b078e715398 · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Self-supervised learning from images with a joint-embedding predictive architecture
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation c9a3fce8-87bf-422a-ba77-fa9d7cacfb40 · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Hallucination of Multimodal Large Language Models: A Survey
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 83cb824a-d8a7-444d-adf8-035a30653275 · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models From colouring-in to pointillism: revisiting semantic segmentation supervision
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c231a48e-3a5c-4215-a6d1-6c23b9219b01 · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Multimodal Autoregressive Pre-training of Large Vision Encoders
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c9d423cd-3234-4f39-bc5c-59a7505f519c · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models The Llama 3 Herd of Models
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 2ce9a543-525d-4b16-b5f1-56ab96fb1fc6 · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation e565e41f-649b-4594-8d4e-23f0bb47a7f3 · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 058f94bb-38e4-4873-8a6f-cdc5645d5cc8 · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models
Reference 9
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b9807f5c-88de-4f85-9435-c86598a4344e · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Visual instruction tuning
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation ba66a892-f2be-4524-a828-e81cf705a085 · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Learning transferable visual models from natural language supervi- sion
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6fc286c2-30da-43c5-af8c-a801e47ba9f7 · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Vision language models are blind
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation a3d89823-2811-4467-8393-ecaefd6298ee · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Am-radio: Agglomerative vision foundation model reduce all domains into one
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 3999e065-746f-4cf2-9d14-73ee2e0163bb · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 013f002a-5d4e-4f4c-ae2d-b02673928a89 · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Eagle: Exploring the design space for multimodal llms with mixture of encoders, 2025
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d6876a80-5e66-4a3c-89eb-4126c3eeee25 · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models An empirical analysis on spatial reason- ing capabilities of large multimodal models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 9efa1615-e701-4c5e-8364-7242022bc0b6 · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Sparkle: Mastering basic spatial capabilities in vision language models elicits gen- eralization to composite spatial reasoning
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18a6c9d6-e108-462e-a853-f4af6e022a58 · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation dbca669c-abaa-4d6c-8a43-a6df314ccbcb · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b5c0014b-a35e-438f-8a7d-38f912c2dc6b · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Is a picture worth a thousand words? delving into spatial reasoning for vi- sion language models
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 2ab09e69-0a89-4155-980e-9eb6e2ac21cc · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Cogvlm: Visual expert for pretrained lan- guage models
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation d4faa040-7b49-45a5-bd45-318db2895f37 · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models MM-LLMs: Recent Advances in MultiModal Large Language Models
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d865d26b-9e31-4d0f-be5e-ab97492a1744 · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Llava-grounding: Grounded visual chat with large multimodal models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation 1401d02a-b2a4-4778-9fd6-99ef7e1098a0 · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bd48bbfe-147f-4335-808d-61f7fae9aff1 · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Unresolved cited work
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
Observation f920a5cb-03e8-4ee9-a0ea-a7d27d415fdb · outbound
Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Unresolved cited work
Reference 2024
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.
No inbound Pith citation observations are available.