Pith. sign in

Paper Citation Record · LEDGER

Hidden in plain sight: VLMs overlook their visual representations

As of 13 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 19 inbound Pith citation observations for arXiv:2506.08008.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08008 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:25:31.975509Z

measured 61 of 61 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-13T06:32:02.005865+00:00

measured 19 of 19 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-06T18:44:45.178051Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:29:29.186840Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d1878cd5-cea2-4902-b674-fde8032d49ac · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Hidden in plain sight: VLMs overlook their visual representations Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.827278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.827278Z digest=sha256:34c13a60f450e3d0f32452e842ba878a506b79e784afedc6358f4a75072f640e

Observation 3da569a7-e35e-43b7-831e-8a15fb0eeb92 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Hidden in plain sight: VLMs overlook their visual representations Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.832536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.832536Z digest=sha256:433ea1733fcb6e8f0d3f49297d063eccf8ad6b9147a597e0f6978c325667a678

Observation 4f5f4e8e-52ec-447f-a42e-25e53966848f · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Hidden in plain sight: VLMs overlook their visual representations OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.836549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.836549Z digest=sha256:f0a7ec5ab301ac18242f54ad52019b7b7b741f23dbfb025243482b883bae60aa

Observation d29bfddc-1992-4023-be6f-1dcf41c3317c · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Hidden in plain sight: VLMs overlook their visual representations Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.840203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.840203Z digest=sha256:dd3901543bcc77afbacc0a6cd95299840c1c8ce1150be33b538312c457a203f9

Observation faf576ad-2880-4a83-927c-45829b2203b0 · outbound

This paper cites Hpatches: A benchmark and evaluation of handcrafted and learned local descriptors.

Hidden in plain sight: VLMs overlook their visual representations Hpatches: A benchmark and evaluation of handcrafted and learned local descriptors

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:32.658742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T05:25:31.843707Z digest=sha256:a394ef5a710b17cf0c6497029e0b9cee33cfd7807cab0f1f545e20f8e27bc551

Observation 580c1030-0ba6-4645-b8a8-c33cea9c1479 · outbound

This paper cites Probing the 3D Awareness of Visual Foundation Models.

Hidden in plain sight: VLMs overlook their visual representations Probing the 3D Awareness of Visual Foundation Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.847196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.847196Z digest=sha256:d10c9f094d3d6d5e3035a66bfc5d07d11aa73160aa5b781fcd428885fd11009a

Observation e26744c1-08d4-4e66-a479-43e0be18273b · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Hidden in plain sight: VLMs overlook their visual representations PaliGemma: A versatile 3B VLM for transfer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.851720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.851720Z digest=sha256:56a90d942fdcdf74471ecc1b8485ce9c3c9b49ee6b4867c068fef326c01f3303

Observation 7cd8f14e-8f94-4bec-a68b-2d43b66e6b0c · outbound

This paper cites Evaluating Multiview Object Consistency in Humans and Image Models.

Hidden in plain sight: VLMs overlook their visual representations Evaluating Multiview Object Consistency in Humans and Image Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.855332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.855332Z digest=sha256:ee39012df8e4c7b5e296c3a58f92a32525fe96cad5be68a885aa5fcc3c8572be

Observation c760c7e7-9e4f-4753-a33c-653eb22cef32 · outbound

This paper cites Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild.

Hidden in plain sight: VLMs overlook their visual representations Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.858960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.858960Z digest=sha256:83ddfb0bed3b8437c8e7b5f224ed74e2e8bbe6846961228e01ab94707021a0a4

Observation bde34bae-c1ab-4412-b835-e865f41fe766 · outbound

This paper cites ShapeNet: An Information-Rich 3D Model Repository.

Hidden in plain sight: VLMs overlook their visual representations ShapeNet: An Information-Rich 3D Model Repository

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.862693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.862693Z digest=sha256:54779af093e798beb9fbddfb3f6d84bd2f665e7314551e1ea8a604635b3d72cd

Observation 6bf1eaae-e437-4f15-9902-44650ed9bd71 · outbound

This paper cites An Empirical Study of Training Self-Supervised Vision Transformers.

Hidden in plain sight: VLMs overlook their visual representations An Empirical Study of Training Self-Supervised Vision Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.866030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.866030Z digest=sha256:717cdfe07bfb7e69fe2658c4f7f3fdab0bdc3b91dfbfaac3e40877141e9705e6

Observation 5dab55f0-ed59-4e5a-8f40-64a1bd16929b · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Hidden in plain sight: VLMs overlook their visual representations Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:32.648354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T05:25:31.870323Z digest=sha256:92347a820b7357f9f6ffe5f7ead8ba9c79e7f77313bd85ce70dfd5fe5cec39dd

Observation f89ad20a-b7b2-4269-8be2-b30b28854f2a · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Hidden in plain sight: VLMs overlook their visual representations Gonzalez, Ion Stoica, and Eric P

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.873622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.873622Z digest=sha256:c1939d2218073be81d3431e833adffd5a1757177b281efbd8d39725a9abbf4a2

Observation b9705a3f-eb6c-4689-b927-f444ef4c5b14 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Hidden in plain sight: VLMs overlook their visual representations Imagenet: A large-scale hierarchical image database

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.877003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.877003Z digest=sha256:c0dafc145e696c6130bdcdb3157927062e45ef9eebaf91446f509949b65fd2c3

Observation 6324917d-df62-4403-96b1-d25a6accdb8f · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Hidden in plain sight: VLMs overlook their visual representations An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.880937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.880937Z digest=sha256:5116debc43929a603bb073204a2638b709582d71c62df8989e793f63146835f2

Observation 37462020-3f3a-46e7-8157-55b8d09095c0 · outbound

This paper cites MouSi: Poly-Visual-Expert Vision-Language Models.

Hidden in plain sight: VLMs overlook their visual representations MouSi: Poly-Visual-Expert Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.884425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.884425Z digest=sha256:68e9d9e9ade9c62dd0924695d4a36ed016439b6f9fa2e711c4b97328ca27fd33

Observation fb1d3bfd-f903-4eff-8b91-cd78b2677d51 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Hidden in plain sight: VLMs overlook their visual representations BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.888488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.888488Z digest=sha256:3b179221385261296dd0f66733d777c0cf760c632f92f3a41fac7338044241d7

Observation 535000f8-1b2d-4c6a-9863-30297af8d5d9 · outbound

This paper cites A Neural Algorithm of Artistic Style.

Hidden in plain sight: VLMs overlook their visual representations A Neural Algorithm of Artistic Style

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.892084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.892084Z digest=sha256:6c47ed4f8f6a300941b3e64d160fe7861e36e97f7ab5d394fa01f4ba8e6a9c3c

Observation 28b990cb-d837-4715-8c96-4b3604a8777f · outbound

This paper cites Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models.

Hidden in plain sight: VLMs overlook their visual representations Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.895470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.895470Z digest=sha256:da03842b9dbb73711097c1b8b491cdf444d29b83c878224ff22c90df5e50cb51

Observation 9200f337-18a2-4632-9222-10fe1cd965fd · outbound

This paper cites The functional correspondence problem.

Hidden in plain sight: VLMs overlook their visual representations The functional correspondence problem

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:32.632004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T05:25:31.899045Z digest=sha256:d790bf62fb82b0099fa4edd3357ab0e42dc2ec1c4cda8e5786ac0e44a46de2ba

Observation b3675a38-0f23-4adb-a3f8-4399078d05e8 · outbound

This paper cites What matters when building vision-language models?.

Hidden in plain sight: VLMs overlook their visual representations What matters when building vision-language models?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.902130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.902130Z digest=sha256:2d72d1ad919038ae26b8c4837ead51168e562abd77b963b496b03c2d6ed6e300

Observation aa320425-5567-4e1b-a1b2-e289f5b01d27 · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

Hidden in plain sight: VLMs overlook their visual representations The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.905842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.905842Z digest=sha256:96df94e53a8a77bf1a6c2115da88874ebeb1e96a018683f474e61a0a5b2e530d

Observation 5150e7b4-13cd-457a-9a7f-9b7b2c84bded · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Hidden in plain sight: VLMs overlook their visual representations Improved Baselines with Visual Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.909053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.909053Z digest=sha256:a137b8ce4de73b3fd35e8fa4ff8178167e0217d7213ddf87973cb8e9f40823f5

Observation 361ab3f8-5750-485e-9c03-72e1fa7965f0 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Hidden in plain sight: VLMs overlook their visual representations Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.912439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.912439Z digest=sha256:a72c8d73964c5925f6b2409f0f6c6a739b9c706225645fcee4173ec5b03efa80

Observation e94772e6-f240-42d1-a7de-8f870d376c8f · outbound

This paper cites Transformer-based neural texture synthesis and style transfer.

Hidden in plain sight: VLMs overlook their visual representations Transformer-based neural texture synthesis and style transfer

Reference 25

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T05:25:32.344244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T05:25:31.915449Z digest=sha256:0b7ad464b5d794c07dcd6da5a4848c93fb9ab6b813f11bda5b9577cc324d731a

Observation b9114a3f-163d-495a-a67c-fd0091df70d4 · outbound

This paper cites SPair-71k: A Large-scale Benchmark for Semantic Correspondence.

Hidden in plain sight: VLMs overlook their visual representations SPair-71k: A Large-scale Benchmark for Semantic Correspondence

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.919369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.919369Z digest=sha256:aa5f3d20bc5b2565fc7bc9d970dce5a13718157dc3f9601bc1ec681146367f58

Observation 44a2ae01-b915-4b82-bab3-8e94fa98192a · outbound

This paper cites Indoor segmentation and support inference from rgbd images.

Hidden in plain sight: VLMs overlook their visual representations Indoor segmentation and support inference from rgbd images

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.922941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.922941Z digest=sha256:e8c1bfe8d28005bf31012300b4e9fae259763eccc308349b059299c5b49562cf

Observation e001612a-9f8c-484f-94b2-ab6402190f12 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Hidden in plain sight: VLMs overlook their visual representations DINOv2: Learning Robust Visual Features without Supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.926174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.926174Z digest=sha256:89ba426ae53f9bcb713d8d6f04595c598c716bc348c34ec270e4a755d411b86e

Observation e89a713b-4577-4ba0-9b57-9d3534bdff1b · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Hidden in plain sight: VLMs overlook their visual representations Learning Transferable Visual Models From Natural Language Supervision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.930340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.930340Z digest=sha256:54a347f4d86743b5c8b3f097f5eea32b0373a4af1601e3c9454ad0cc3287e3da

Observation a61c7834-0bfd-4419-8d62-6bfe83bf8bad · outbound

This paper cites Vision Transformers for Dense Prediction.

Hidden in plain sight: VLMs overlook their visual representations Vision Transformers for Dense Prediction

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.933811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.933811Z digest=sha256:39aa4bca417d147623ee3349d6fd97244f69e46cc10c71834c5a5d85847fe3a5

Observation 1cb2a407-73a5-493f-bfe7-de063c93e6c1 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

Hidden in plain sight: VLMs overlook their visual representations Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.936899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.936899Z digest=sha256:269d97d2cd2ef7940cf7b1216890317c2b4f7513b3a6979d07ddea3c1daa6e56

Observation 96bd96aa-b993-4134-9747-392bb262feb0 · outbound

This paper cites How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers.

Hidden in plain sight: VLMs overlook their visual representations How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.940685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.940685Z digest=sha256:70c9f7c494095411a3ca186c6fe2bab4443726d603536d7f8df6d4b5d908d049

Observation 944a45ef-63b3-4ef9-b4f7-15b93ad1ce04 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Hidden in plain sight: VLMs overlook their visual representations Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.944023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.944023Z digest=sha256:9d3c93707037be9c9def128273381cb7b691a33b721517393a09a5cd54a7804e

Observation 6972fb65-294c-42bc-b23c-9462b1bfbb11 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Hidden in plain sight: VLMs overlook their visual representations Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.947615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.947615Z digest=sha256:a5f83337815439c76b46dc11086e000543e5ccec8c7d57900231107c8f3f0c16

Observation 237f0622-530e-4866-bb3b-e867106bf4de · outbound

This paper cites Disn: Deep implicit surface network for high-quality single-view 3d reconstruction.

Hidden in plain sight: VLMs overlook their visual representations Disn: Deep implicit surface network for high-quality single-view 3d reconstruction

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:32.604495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-08-07T05:25:31.950731Z digest=sha256:a95b1fb3fadc571fd28e2c41db94b94e62dd38f782ae336aff3f7d64a9476233

Observation 8b48f5f0-46a2-4ef5-aaaf-6c5027b37b2d · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

Hidden in plain sight: VLMs overlook their visual representations MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.954005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.954005Z digest=sha256:807e2f7132d6eeb92e3a01b203c7bbb28e5803b8267d720b84896462ee9a4cb7

Observation aec4e2d1-989f-45d6-b10f-c0998f888da2 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Hidden in plain sight: VLMs overlook their visual representations Sigmoid Loss for Language Image Pre-Training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.957341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.957341Z digest=sha256:f32962f6bba9e4f694ba8ccbfbac22909f6132f27487f0dc946955d91ef9ca47

Observation f66242d9-6b7d-4b1b-bf6c-4d27d675e400 · outbound

This paper cites A General Protocol to Probe Large Vision Models for 3D Physical Understanding.

Hidden in plain sight: VLMs overlook their visual representations A General Protocol to Probe Large Vision Models for 3D Physical Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.960678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.960678Z digest=sha256:46a8ec157079440429f2d82b8e4434c2c6be65b730b60c555a6f13944c9fe311

Observation 097b52e8-2d6e-47ad-9065-505b1eeaa5d6 · outbound

This paper cites write newline.

Hidden in plain sight: VLMs overlook their visual representations write newline

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.964053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.964053Z digest=sha256:736eee5bf45e5ac4cdfd38482465a122ba418bd770dc4aa34d823db3a5d403f1

Observation 5aeb6766-572d-484d-9d53-34305de2ff79 · outbound

This paper cites @esa (Ref.

Hidden in plain sight: VLMs overlook their visual representations @esa (Ref

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.967916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.967916Z digest=sha256:71e73430747d707f805722917f1b9cc5286d78c6da5820c5e855f3f5ba638c25

Observation 8edfdc6e-430b-43eb-b4c3-34f7e275f8d4 · outbound

This paper cites an unresolved cited work.

Hidden in plain sight: VLMs overlook their visual representations Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.971733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.971733Z digest=sha256:319987e072b0f9021a940a09da1abf3b8154152717c4e997b7657ab646c826c6

Observation f879e721-808e-4138-b4c2-9b7742b2229a · outbound

This paper cites A, B, C, D.

Hidden in plain sight: VLMs overlook their visual representations A, B, C, D

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.975509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.975509Z digest=sha256:6c5d66e25623384749116787fa84b8f47834a906deaa952bfe03a612506e0ad7

Pith citing papers

Observation db4f2ae1-7184-47fd-8352-468f01ff091c · inbound

PictOBI-20k: Unveiling Large Multimodal Models in Visual Decipherment for Pictographic Oracle Bone Characters cites this paper.

PictOBI-20k: Unveiling Large Multimodal Models in Visual Decipherment for Pictographic Oracle Bone Characters Hidden in plain sight: VLMs overlook their visual representations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T05:02:23.812923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:02:23.812923Z digest=sha256:0580db9965a58c94ec3b75a5824948fd078239fbeeee3f317e447ea8dc19a872

Observation 3a901a8f-b164-4932-a899-d8b48b76f59b · inbound

SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning cites this paper.

SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning Hidden in plain sight: VLMs overlook their visual representations

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:24:21.380035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T20:24:02.748854Z digest=sha256:ad41858f064a0a89f327edfdf8930429e8f0b97550bf1f9982a9ba1ff5ff4fce

Observation 906c870a-fb4f-47a1-bc78-34f6579d5ef1 · inbound

Mull-Tokens: Modality-Agnostic Latent Thinking cites this paper.

Mull-Tokens: Modality-Agnostic Latent Thinking Hidden in plain sight: VLMs overlook their visual representations

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:58:38.582363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-16T22:57:04.802871Z digest=sha256:8e65ab1f460223923e5c293351a73124c5641ef5d6f3ad0c8775c285d7e941ee

Observation 0ece36ae-f997-4c5f-aced-1b408dbee82a · inbound

Egocentric Bias in Vision-Language Models cites this paper.

Egocentric Bias in Vision-Language Models Hidden in plain sight: VLMs overlook their visual representations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:42.981300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:42.981300Z digest=sha256:65cc7ef85d6733b1780b3461479918f54b546794d554d880161b6d989c13ff75

Observation e02c4924-6ba0-4eec-9be2-0cddcca0ab73 · inbound

Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation cites this paper.

Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation Hidden in plain sight: VLMs overlook their visual representations

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:20:07.008117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-21T12:15:57.019720Z digest=sha256:30745687c798b8a6e4a3c91c9b534126e47f3c43c1f4653315926a661e81fbef

Observation 843cc588-bcef-45b4-ad02-e53dfa76a3d0 · inbound

Vision Language Models Cannot Reason About Physical Transformation cites this paper.

Vision Language Models Cannot Reason About Physical Transformation Hidden in plain sight: VLMs overlook their visual representations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-15T13:27:51.848177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:27:51.848177Z digest=sha256:976ae7c564c005905dcc4770eb20e1d6de3cd698bced7d22092ae33de4074c79

Observation 4825b433-659f-4773-8bee-187497140f64 · inbound

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors cites this paper.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors Hidden in plain sight: VLMs overlook their visual representations

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T21:58:19.843334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:11b268580974ffb0f5a712b7dde49f02c773c5b2f020c0bd80cdebac1df7f466

Observation e41c6df5-c7aa-4fd7-9275-eec59f7b92e8 · inbound

Watch Before You Answer: Learning from Visually Grounded Post-Training cites this paper.

Watch Before You Answer: Learning from Visually Grounded Post-Training Hidden in plain sight: VLMs overlook their visual representations

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:20:46.803606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T20:01:13.305374Z digest=sha256:f602c5f8cd0c8e22870fec9efd51fa8291e2ba4590b3c82989b61e4296c36ecc

Observation deee1951-098a-4113-85f5-857b9badc27d · inbound

Lost in the Hype: Revealing and Dissecting the Performance Degradation of Medical Multimodal Large Language Models in Image Classification cites this paper.

Lost in the Hype: Revealing and Dissecting the Performance Degradation of Medical Multimodal Large Language Models in Image Classification Hidden in plain sight: VLMs overlook their visual representations

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:21:00.893146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T18:11:13.523021Z digest=sha256:158c3a2d9c5e5084972e93d668a256254dd35b471a650edfb4cd0cc951adb14f

Observation 934e2904-d7e1-4c89-978d-3a8e55ac61cf · inbound

Do Vision Language Models Need to Process Image Tokens? cites this paper.

Do Vision Language Models Need to Process Image Tokens? Hidden in plain sight: VLMs overlook their visual representations

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:35:50.692738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T18:26:37.371488Z digest=sha256:f8873f0ecbb23170d4130810187cadd8bbf1077d41bb981b4119dab73036251b

Observation 2c8aec45-5da4-4685-9d43-3c26b0e35eca · inbound

Boosting Visual Instruction Tuning with Self-Supervised Guidance cites this paper.

Boosting Visual Instruction Tuning with Self-Supervised Guidance Hidden in plain sight: VLMs overlook their visual representations

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:50:58.634642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T16:27:52.208827Z digest=sha256:3f9ca289611bd8f0df26320b205b3d4ddbcac55fa45215bf9dd399e02df5e59d

Observation df2215f2-cab2-497c-97a8-f12aad1d624d · inbound

The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models cites this paper.

The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models Hidden in plain sight: VLMs overlook their visual representations

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:35:26.811566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-10T13:25:27.762910Z digest=sha256:9eb558f5389707a07402cc875a263e9fdcfa65d1c97e049db7196d612f1efddb

Observation 575432cf-dadc-494c-807f-1a53dc2b3433 · inbound

Do multimodal models imagine electric sheep? cites this paper.

Do multimodal models imagine electric sheep? Hidden in plain sight: VLMs overlook their visual representations

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.524703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-12T03:24:01.933339Z digest=sha256:05915a02872acbda212544e9afaf5e0d1c43aef240da454fa37b70e31ff36242

Observation 9dc668a8-f105-4526-8b89-61ae92877cf9 · inbound

A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning cites this paper.

A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning Hidden in plain sight: VLMs overlook their visual representations

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:17:51.063386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-14T19:14:03.809826Z digest=sha256:36cab37b59ab1cfb164738cdc5e465930985910c91f80bb37e889158b7b59e8d

Observation dc64514a-24f0-456a-b219-b2dbf321a2ba · inbound

A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning cites this paper.

A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning Hidden in plain sight: VLMs overlook their visual representations

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:09:07.103332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-05-20T22:07:35.021986Z digest=sha256:2abbdbcdc83e3b0c29d9cc69f03bf3d9fed448b4a72bb14333fb95e688caa187

Observation b957b92a-4bcf-4ae6-8c5e-30ce662dcb60 · inbound

Diagnosing Visual Ignorance in Vision-Language Models cites this paper.

Diagnosing Visual Ignorance in Vision-Language Models Hidden in plain sight: VLMs overlook their visual representations

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:17:09.193506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=arxiv_source observed=2026-06-27T22:51:25.269986Z digest=sha256:2ae557ab3acca9226173ba6fdf91217e8d469ed47ea41470534743f2131c238e

Observation 43114add-9cbc-41f2-99ff-81fde306c970 · inbound

The Hidden Evolution of Disguised Visual Context inside the VLM cites this paper.

The Hidden Evolution of Disguised Visual Context inside the VLM Hidden in plain sight: VLMs overlook their visual representations

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:29:29.190447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-13T06:32:02.005865+00:00.

source=pdf_text observed=2026-06-26T18:08:56.044278Z digest=sha256:f6b8d651408bae2561a190cb205e365d35689fd925aad1b03c69e3a7fac6057a

Observation 4ab637ef-2e2c-4309-8082-436d893e5348 · inbound

Visual Access Boundaries in Vision-Language Model Reasoning cites this paper.

Visual Access Boundaries in Vision-Language Model Reasoning Hidden in plain sight: VLMs overlook their visual representations

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T06:24:27.325903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:24:27.325903Z digest=sha256:75354163a0f9ed9595647ef06d0dd90a00472a425049d81cec829f5379faf329

Observation 3fd8021e-4afd-4b44-89b7-8baeb11f729c · inbound

Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO cites this paper.

Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO Hidden in plain sight: VLMs overlook their visual representations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:44:45.178051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:44:45.178051Z digest=sha256:148703eb5128287b81db73f1380568055f12177624391f7d8942506055a3bbbc