Pith. sign in

Paper Citation Record · LEDGER

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models

As of 16 August 2026, this Paper Citation Record lists 26 of 26 outbound references and 0 inbound Pith citation observations for arXiv:2505.05626.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2505.05626 v3

Coverage vector

measured 26 of 26 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T23:06:49.290784Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-16T06:30:59.297886+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

26 of 26 outbound references displayed

  • verified exact0
  • verified fuzzy13
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 3dace842-e939-43af-b761-3563d2ff32e4 · outbound

This paper cites an unresolved cited work.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:06:49.692845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.185329Z digest=sha256:2bba832123cc9ed1466b05ffff7b02bed4a1a00db78cfb3d2b92ce97d6df202f

Observation 0d8090c1-07b4-4f97-ad1d-4b078e715398 · outbound

This paper cites Self-supervised learning from images with a joint-embedding predictive architecture.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Self-supervised learning from images with a joint-embedding predictive architecture

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.679539Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.190221Z digest=sha256:cd831ddaf0ff026f17a4d5a6b6300f3d58582f0803fbc20854bd28b2da670686

Observation c9a3fce8-87bf-422a-ba77-fa9d7cacfb40 · outbound

This paper cites Hallucination of Multimodal Large Language Models: A Survey.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Hallucination of Multimodal Large Language Models: A Survey

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.194642Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.194642Z digest=sha256:d044fbe0385217598d0866cb8228c6f4fc7d4059ad2ef8883c6ac11bd2d20205

Observation 83cb824a-d8a7-444d-adf8-035a30653275 · outbound

This paper cites From colouring-in to pointillism: revisiting semantic segmentation supervision.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models From colouring-in to pointillism: revisiting semantic segmentation supervision

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.199068Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.199068Z digest=sha256:d501ca2b5a3d8c87304f9aa9a67d23b87a3564ff1656844418e8e9dd11322541

Observation c231a48e-3a5c-4215-a6d1-6c23b9219b01 · outbound

This paper cites Multimodal Autoregressive Pre-training of Large Vision Encoders.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Multimodal Autoregressive Pre-training of Large Vision Encoders

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.204156Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.204156Z digest=sha256:1aa8fa9a67344d47ef37217ce1b6c5cab051eb6366354e875ff2e4c934595307

Observation c9d423cd-3234-4f39-bc5c-59a7505f519c · outbound

This paper cites The Llama 3 Herd of Models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models The Llama 3 Herd of Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.208740Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.208740Z digest=sha256:07a02d2ada278710f0e5b74ae04e444c384d625880447fd0d9e298e96d3bf814

Observation 2ce9a543-525d-4b16-b5f1-56ab96fb1fc6 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.666781Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.214294Z digest=sha256:b28566b9d9a6fe41c2d3d38cc50bb661f105b1f5eeb73360600f53c226ea1613

Observation e565e41f-649b-4594-8d4e-23f0bb47a7f3 · outbound

This paper cites The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models The open images dataset v4: Unified image classification, object detection, and visual relationship detection at scale

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.654002Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.218511Z digest=sha256:f60d31f15de6be2587e391fadb6091ce8a352ae00322a9aa4162aae3a53d4e68

Observation 058f94bb-38e4-4873-8a6f-cdc5645d5cc8 · outbound

This paper cites Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Mixture-of-Transformers: A Sparse and Scalable Architecture for Multi-Modal Foundation Models

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.222650Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.222650Z digest=sha256:88318bb2f3fad8523b686dd8f879f51ad5cac4016fceb76b2b653e6455f2b7b8

Observation b9807f5c-88de-4f85-9435-c86598a4344e · outbound

This paper cites Visual instruction tuning.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Visual instruction tuning

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.640996Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.226922Z digest=sha256:2e0d5edf07206b3627ffa8b8581b25eb26e1ba49e984c9a99d72f69281e87fdf

Observation ba66a892-f2be-4524-a828-e81cf705a085 · outbound

This paper cites Learning transferable visual models from natural language supervi- sion.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Learning transferable visual models from natural language supervi- sion

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.230819Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.230819Z digest=sha256:777843fd046bc764a30fef839e2b041619a7e46aeada31988e6ab2cb13fc12bf

Observation 6fc286c2-30da-43c5-af8c-a801e47ba9f7 · outbound

This paper cites Vision language models are blind.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Vision language models are blind

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.619327Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.234940Z digest=sha256:cdedb75ae76d2dc378443540253593848ac96b5bba73a15e57b880649e1421d6

Observation a3d89823-2811-4467-8393-ecaefd6298ee · outbound

This paper cites Am-radio: Agglomerative vision foundation model reduce all domains into one.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Am-radio: Agglomerative vision foundation model reduce all domains into one

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.605236Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.238981Z digest=sha256:bfbcb49274f53958988eb4ed436cffc928f207c24a3d648ca737d2e761641292

Observation 3999e065-746f-4cf2-9d14-73ee2e0163bb · outbound

This paper cites Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Anwer, Eric Xing, Ming-Hsuan Yang, and Fahad S

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.591033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.242786Z digest=sha256:0287eec056eab418d6fcd0b39b17fd30a27a3cabf93cee5dbcf5c9db23c34df5

Observation 013f002a-5d4e-4f4c-ae2d-b02673928a89 · outbound

This paper cites Eagle: Exploring the design space for multimodal llms with mixture of encoders, 2025.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Eagle: Exploring the design space for multimodal llms with mixture of encoders, 2025

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.578106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.246827Z digest=sha256:2577fd664d2a38d4a31c8b457a70c6847c1459d6bd0a3cf7be3389d5d7698284

Observation d6876a80-5e66-4a3c-89eb-4126c3eeee25 · outbound

This paper cites An empirical analysis on spatial reason- ing capabilities of large multimodal models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models An empirical analysis on spatial reason- ing capabilities of large multimodal models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.565288Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.250843Z digest=sha256:1ddc1da638ebd281b0dfeacf8267084e423c7c55cc854a661e8880cf0cad4a3e

Observation 9efa1615-e701-4c5e-8364-7242022bc0b6 · outbound

This paper cites Sparkle: Mastering basic spatial capabilities in vision language models elicits gen- eralization to composite spatial reasoning.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Sparkle: Mastering basic spatial capabilities in vision language models elicits gen- eralization to composite spatial reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.254574Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.254574Z digest=sha256:4acb7db7c0d676b4ca8f0a31a86344d0f3e7d8c76b040fc477049d4de7834dea

Observation 18a6c9d6-e108-462e-a853-f4af6e022a58 · outbound

This paper cites Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Cambrian-1: A fully open, vision-centric explo- ration of multimodal llms

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.552250Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.258487Z digest=sha256:343360911f12d0cf9e5f5ba19fe5fdaf250b2269db71f732ba771eed1a26c6e9

Observation dbca669c-abaa-4d6c-8a43-a6df314ccbcb · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.262424Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.262424Z digest=sha256:f7da9e910897483818c0308323136744d55453f5efbd7172026f06ac04915045

Observation b5c0014b-a35e-438f-8a7d-38f912c2dc6b · outbound

This paper cites Is a picture worth a thousand words? delving into spatial reasoning for vi- sion language models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Is a picture worth a thousand words? delving into spatial reasoning for vi- sion language models

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.530294Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.266213Z digest=sha256:85ca855d183daa585a06b9a92472ddad8a4bb1ee3e78fb4bc07ced7f12f67cd8

Observation 2ab09e69-0a89-4155-980e-9eb6e2ac21cc · outbound

This paper cites Cogvlm: Visual expert for pretrained lan- guage models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Cogvlm: Visual expert for pretrained lan- guage models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.516505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.270221Z digest=sha256:1150187af591ebe2c5c77d7be5a4a51048941474cb4264838382516cdb08b02d

Observation d4faa040-7b49-45a5-bd45-318db2895f37 · outbound

This paper cites MM-LLMs: Recent Advances in MultiModal Large Language Models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models MM-LLMs: Recent Advances in MultiModal Large Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.274106Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.274106Z digest=sha256:1cd427651ea39979a70c5ea5b7250ecb523959142d2131ce3e4a5a9040eda1b7

Observation d865d26b-9e31-4d0f-be5e-ab97492a1744 · outbound

This paper cites Llava-grounding: Grounded visual chat with large multimodal models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Llava-grounding: Grounded visual chat with large multimodal models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T23:06:49.503755Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.278600Z digest=sha256:4168b202eddd36920de45116648ba5ce49567e1a765cee71ba08dd16666cad31

Observation 1401d02a-b2a4-4778-9fd6-99ef7e1098a0 · outbound

This paper cites Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Ferret-v2: An Improved Baseline for Referring and Grounding with Large Language Models

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-15T23:06:49.282590Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T23:06:49.282590Z digest=sha256:746af3292c886a125120aaa37a402adb2dc5aa7ab471e288b88b9a6959ead599

Observation bd48bbfe-147f-4335-808d-61f7fae9aff1 · outbound

This paper cites an unresolved cited work.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Unresolved cited work

Reference 26

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:06:49.476853Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.290784Z digest=sha256:0fe9118589bc758ec5bd7284c1926a62969fb075d255adeef4aa162635746341

Observation f920a5cb-03e8-4ee9-a0ea-a7d27d415fdb · outbound

This paper cites an unresolved cited work.

Perceiving Beyond Language Priors: Enhancing Visual Comprehension and Attention in Multimodal Models Unresolved cited work

Reference 2024

Resolution
unresolved
raw_fallback, observed 2026-08-15T23:06:49.490235Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-16T06:30:59.297886+00:00.

source=pdf_text observed=2026-08-15T23:06:49.286769Z digest=sha256:d7b571d4770195439370dccbca635cee8e69cace48f5e6549106a9a621349bb8

Pith citing papers

No inbound Pith citation observations are available.