Pith. sign in

Paper Citation Record · LEDGER

Hidden in plain sight: VLMs overlook their visual representations

As of 22 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 20 inbound Pith citation observations for arXiv:2506.08008.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2506.08008 v1

Coverage vector

measured 42 of 42 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-07T05:25:31.975509Z

measured 62 of 62 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 20 of 20 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-15T15:10:13.066770Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-07-04T03:29:29.186840Z

Reference resolution

42 of 42 outbound references displayed

  • verified exact0
  • verified fuzzy4
  • unresolved37
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d1878cd5-cea2-4902-b674-fde8032d49ac · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

Hidden in plain sight: VLMs overlook their visual representations Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.827278Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.827278Z digest=sha256:35dcc31dc96f956e501f7be3bf00bf544b4aa90684aae1796e057356bc53f641

Observation 3da569a7-e35e-43b7-831e-8a15fb0eeb92 · outbound

This paper cites Flamingo: a Visual Language Model for Few-Shot Learning.

Hidden in plain sight: VLMs overlook their visual representations Flamingo: a Visual Language Model for Few-Shot Learning

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.832536Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.832536Z digest=sha256:3b1d3d4fb939bf3ba5796365edb484ebf719c82948ea4053cf951aead3362376

Observation 4f5f4e8e-52ec-447f-a42e-25e53966848f · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

Hidden in plain sight: VLMs overlook their visual representations OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.836549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.836549Z digest=sha256:7b384ae4bd95dc46941707caa2243aac6680f29a1fb8165edc18b1158376fd72

Observation d29bfddc-1992-4023-be6f-1dcf41c3317c · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

Hidden in plain sight: VLMs overlook their visual representations Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.840203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.840203Z digest=sha256:1e68413894f2ef2e6a1bae30b1f1bf0318dec2b2caaf79b0e90019e1c0681685

Observation faf576ad-2880-4a83-927c-45829b2203b0 · outbound

This paper cites Hpatches: A benchmark and evaluation of handcrafted and learned local descriptors.

Hidden in plain sight: VLMs overlook their visual representations Hpatches: A benchmark and evaluation of handcrafted and learned local descriptors

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:32.658742Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T05:25:31.843707Z digest=sha256:a35732b5938f368a2969be1ceba5f0a84bfecfb3d9f55be0fc581f5a9892eecd

Observation 580c1030-0ba6-4645-b8a8-c33cea9c1479 · outbound

This paper cites Probing the 3D Awareness of Visual Foundation Models.

Hidden in plain sight: VLMs overlook their visual representations Probing the 3D Awareness of Visual Foundation Models

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.847196Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.847196Z digest=sha256:a920d8f4f6e5de0965b2341326ba0a3d3a2a9c7658f1b94e073c006da06eeed3

Observation e26744c1-08d4-4e66-a479-43e0be18273b · outbound

This paper cites PaliGemma: A versatile 3B VLM for transfer.

Hidden in plain sight: VLMs overlook their visual representations PaliGemma: A versatile 3B VLM for transfer

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.851720Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.851720Z digest=sha256:f4d9e496103a41a22f0e7a66e1b9f724a98a0c68535e5dfa1c60ca3fd38da932

Observation 7cd8f14e-8f94-4bec-a68b-2d43b66e6b0c · outbound

This paper cites Evaluating Multiview Object Consistency in Humans and Image Models.

Hidden in plain sight: VLMs overlook their visual representations Evaluating Multiview Object Consistency in Humans and Image Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.855332Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.855332Z digest=sha256:b9cacf003f1f071a48f6e9bbc981715e8118903dc9562e4e37c1c4ac0dcbc4f1

Observation c760c7e7-9e4f-4753-a33c-653eb22cef32 · outbound

This paper cites Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild.

Hidden in plain sight: VLMs overlook their visual representations Omni3D: A Large Benchmark and Model for 3D Object Detection in the Wild

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.858960Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.858960Z digest=sha256:f738d396834c2c86e64387508191059af8fdd6bc5700e87a56751d7d98cb992d

Observation bde34bae-c1ab-4412-b835-e865f41fe766 · outbound

This paper cites ShapeNet: An Information-Rich 3D Model Repository.

Hidden in plain sight: VLMs overlook their visual representations ShapeNet: An Information-Rich 3D Model Repository

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.862693Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.862693Z digest=sha256:e9f08bfafb9f8a7dd6ff5975fc317ace9c722ce0363df7a75c7a3f62561a2890

Observation 6bf1eaae-e437-4f15-9902-44650ed9bd71 · outbound

This paper cites An Empirical Study of Training Self-Supervised Vision Transformers.

Hidden in plain sight: VLMs overlook their visual representations An Empirical Study of Training Self-Supervised Vision Transformers

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.866030Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.866030Z digest=sha256:f7e1a14faa2218a8b6afeb546f1dfb954e4c2e2bb77d93289c4710860a0777da

Observation 5dab55f0-ed59-4e5a-8f40-64a1bd16929b · outbound

This paper cites Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks.

Hidden in plain sight: VLMs overlook their visual representations Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:32.648354Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T05:25:31.870323Z digest=sha256:54aa957683666a254ba03cc1db2dfd932e07259afe0be357699cf086b770a41d

Observation f89ad20a-b7b2-4269-8be2-b30b28854f2a · outbound

This paper cites Gonzalez, Ion Stoica, and Eric P.

Hidden in plain sight: VLMs overlook their visual representations Gonzalez, Ion Stoica, and Eric P

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.873622Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.873622Z digest=sha256:331ea1eb4f115fbc1bf5a19636508a74ee3c6d0a35cf1e2c74cbc15cc4e0b81d

Observation b9705a3f-eb6c-4689-b927-f444ef4c5b14 · outbound

This paper cites Imagenet: A large-scale hierarchical image database.

Hidden in plain sight: VLMs overlook their visual representations Imagenet: A large-scale hierarchical image database

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.877003Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.877003Z digest=sha256:414c8126247958125f97457301c79c969139b23131bb582519bd2dbd9ea46e82

Observation 6324917d-df62-4403-96b1-d25a6accdb8f · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

Hidden in plain sight: VLMs overlook their visual representations An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.880937Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.880937Z digest=sha256:666b069ed8d7c04893d2560adce4eb14442469b4f12b5347584dd25dcdc27df6

Observation 37462020-3f3a-46e7-8157-55b8d09095c0 · outbound

This paper cites MouSi: Poly-Visual-Expert Vision-Language Models.

Hidden in plain sight: VLMs overlook their visual representations MouSi: Poly-Visual-Expert Vision-Language Models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.884425Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.884425Z digest=sha256:5c889bc1c6b6185e5456122c0d77f7af929fcfb3902985fbdcb48dd118d49e3d

Observation fb1d3bfd-f903-4eff-8b91-cd78b2677d51 · outbound

This paper cites BLINK: Multimodal Large Language Models Can See but Not Perceive.

Hidden in plain sight: VLMs overlook their visual representations BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.888488Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.888488Z digest=sha256:0f7f86f71fd2399ae9b6d6a1c6652d0fa06804ddbbbe1e622777886412de1c29

Observation 535000f8-1b2d-4c6a-9863-30297af8d5d9 · outbound

This paper cites A Neural Algorithm of Artistic Style.

Hidden in plain sight: VLMs overlook their visual representations A Neural Algorithm of Artistic Style

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.892084Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.892084Z digest=sha256:e3c65d9d4ae919ddd59041ab070bace075875a0417d37f08ae8f047836c421aa

Observation 28b990cb-d837-4715-8c96-4b3604a8777f · outbound

This paper cites Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models.

Hidden in plain sight: VLMs overlook their visual representations Prismatic VLMs: Investigating the Design Space of Visually-Conditioned Language Models

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.895470Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.895470Z digest=sha256:6261b38fe7198cad2d7655d938af6816348b9fe7b5803a18e0820e7775d7805e

Observation 9200f337-18a2-4632-9222-10fe1cd965fd · outbound

This paper cites The functional correspondence problem.

Hidden in plain sight: VLMs overlook their visual representations The functional correspondence problem

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:32.632004Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T05:25:31.899045Z digest=sha256:b9d46df81fce380f0d8265e1756ceed6a37fa9d8e124999408b191f637688b38

Observation b3675a38-0f23-4adb-a3f8-4399078d05e8 · outbound

This paper cites What matters when building vision-language models?.

Hidden in plain sight: VLMs overlook their visual representations What matters when building vision-language models?

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.902130Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.902130Z digest=sha256:cf9bbe52a6f9b5c830987720a5170bfcf42e433f967dba605bca3c1364e71c11

Observation aa320425-5567-4e1b-a1b2-e289f5b01d27 · outbound

This paper cites The Power of Scale for Parameter-Efficient Prompt Tuning.

Hidden in plain sight: VLMs overlook their visual representations The Power of Scale for Parameter-Efficient Prompt Tuning

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.905842Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.905842Z digest=sha256:f960c381cf72e2f31fe1f0bc559970587c48aef4607fb730dfa84c9d3da48c7f

Observation 5150e7b4-13cd-457a-9a7f-9b7b2c84bded · outbound

This paper cites Improved Baselines with Visual Instruction Tuning.

Hidden in plain sight: VLMs overlook their visual representations Improved Baselines with Visual Instruction Tuning

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.909053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.909053Z digest=sha256:93513d9c62ccad3b53bf87e71952c770517533422e46c9d5747466e6d73477d4

Observation 361ab3f8-5750-485e-9c03-72e1fa7965f0 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, January 2024.

Hidden in plain sight: VLMs overlook their visual representations Llava-next: Improved reasoning, ocr, and world knowledge, January 2024

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.912439Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.912439Z digest=sha256:619d9255d5624994b6849d4b55c90ee92aba716f018484dc5542249c23c92e03

Observation e94772e6-f240-42d1-a7de-8f870d376c8f · outbound

This paper cites Transformer-based neural texture synthesis and style transfer.

Hidden in plain sight: VLMs overlook their visual representations Transformer-based neural texture synthesis and style transfer

Reference 25

Resolution
metadata mismatch
raw_fallback, observed 2026-08-07T05:25:32.344244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T05:25:31.915449Z digest=sha256:d9af62c59db8a4fdd173d728722499ea2387e1cd6c40d54cc824cb5318104fb3

Observation b9114a3f-163d-495a-a67c-fd0091df70d4 · outbound

This paper cites SPair-71k: A Large-scale Benchmark for Semantic Correspondence.

Hidden in plain sight: VLMs overlook their visual representations SPair-71k: A Large-scale Benchmark for Semantic Correspondence

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.919369Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.919369Z digest=sha256:6e4d117d2deaca234a54edec5829def2de9c3557a9e4ad9f1e48acb63a148767

Observation 44a2ae01-b915-4b82-bab3-8e94fa98192a · outbound

This paper cites Indoor segmentation and support inference from rgbd images.

Hidden in plain sight: VLMs overlook their visual representations Indoor segmentation and support inference from rgbd images

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.922941Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.922941Z digest=sha256:fd76ecf399b7b3801e1d557968f14f4d511b453b8d4e4b0009d5ae3bf24d77f0

Observation e001612a-9f8c-484f-94b2-ab6402190f12 · outbound

This paper cites DINOv2: Learning Robust Visual Features without Supervision.

Hidden in plain sight: VLMs overlook their visual representations DINOv2: Learning Robust Visual Features without Supervision

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.926174Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.926174Z digest=sha256:10a0c3592ab14c158f4a82fd766faa26c58b691037f4c3ba76bf3da469e6790a

Observation e89a713b-4577-4ba0-9b57-9d3534bdff1b · outbound

This paper cites Learning Transferable Visual Models From Natural Language Supervision.

Hidden in plain sight: VLMs overlook their visual representations Learning Transferable Visual Models From Natural Language Supervision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.930340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.930340Z digest=sha256:69e4018de84ffe440b1aa506fc0ee63a82dce72fc61471f55197a24a2ec56ece

Observation a61c7834-0bfd-4419-8d62-6bfe83bf8bad · outbound

This paper cites Vision Transformers for Dense Prediction.

Hidden in plain sight: VLMs overlook their visual representations Vision Transformers for Dense Prediction

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.933811Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.933811Z digest=sha256:e57168bd98badc6ac831ff853b0c894478ae410d2331ba16c549ede72c70b003

Observation 1cb2a407-73a5-493f-bfe7-de063c93e6c1 · outbound

This paper cites Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders.

Hidden in plain sight: VLMs overlook their visual representations Eagle: Exploring The Design Space for Multimodal LLMs with Mixture of Encoders

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.936899Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.936899Z digest=sha256:0bb17442068a60f681938893b6471209bc8632ec50f1906f8c16bf74d315745b

Observation 96bd96aa-b993-4134-9747-392bb262feb0 · outbound

This paper cites How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers.

Hidden in plain sight: VLMs overlook their visual representations How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.940685Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.940685Z digest=sha256:db4239c97a154250fbba3acd808f7b8b730ccd7b22fba551c5d19528b57064c9

Observation 944a45ef-63b3-4ef9-b4f7-15b93ad1ce04 · outbound

This paper cites Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs.

Hidden in plain sight: VLMs overlook their visual representations Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.944023Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.944023Z digest=sha256:a1ee2cb9fcdc84149b8866cdaef67115c8ffa093f8b0c3e33b1adca7debca818

Observation 6972fb65-294c-42bc-b23c-9462b1bfbb11 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Hidden in plain sight: VLMs overlook their visual representations Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.947615Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.947615Z digest=sha256:2d7dc9a8a4842d7f8e6cee50c59c18f17ab6f27eeae4ea7e5e474264d80113ed

Observation 237f0622-530e-4866-bb3b-e867106bf4de · outbound

This paper cites Disn: Deep implicit surface network for high-quality single-view 3d reconstruction.

Hidden in plain sight: VLMs overlook their visual representations Disn: Deep implicit surface network for high-quality single-view 3d reconstruction

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-07T05:25:32.604495Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-08-07T05:25:31.950731Z digest=sha256:db535cb820eb15a0ff8c6871c2e24d7783a8d0e4c3a5945e93475e7a34609437

Observation 8b48f5f0-46a2-4ef5-aaaf-6c5027b37b2d · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

Hidden in plain sight: VLMs overlook their visual representations MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.954005Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.954005Z digest=sha256:86297f17c4910e59ee9cf4d62badf6802934f598a93ab13c44f4e0850ea2d885

Observation aec4e2d1-989f-45d6-b10f-c0998f888da2 · outbound

This paper cites Sigmoid Loss for Language Image Pre-Training.

Hidden in plain sight: VLMs overlook their visual representations Sigmoid Loss for Language Image Pre-Training

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.957341Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.957341Z digest=sha256:bc7a44832cc23079a1850a38f5889585605b60bb2e7a489965a9839db460c9b4

Observation f66242d9-6b7d-4b1b-bf6c-4d27d675e400 · outbound

This paper cites A General Protocol to Probe Large Vision Models for 3D Physical Understanding.

Hidden in plain sight: VLMs overlook their visual representations A General Protocol to Probe Large Vision Models for 3D Physical Understanding

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.960678Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.960678Z digest=sha256:8c061c51d41564f5befd6aefef106bba3e4aba670d224cce8bf18c8e1e466937

Observation 097b52e8-2d6e-47ad-9065-505b1eeaa5d6 · outbound

This paper cites write newline.

Hidden in plain sight: VLMs overlook their visual representations write newline

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.964053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.964053Z digest=sha256:096f28b303edd6b1887090d9b231beabe676461b525dfddaf9bd932ce1d35c8c

Observation 5aeb6766-572d-484d-9d53-34305de2ff79 · outbound

This paper cites @esa (Ref.

Hidden in plain sight: VLMs overlook their visual representations @esa (Ref

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.967916Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.967916Z digest=sha256:7a5cbc8dbab2310a6d72196ca5240f5f34492f7a14a531f1919d99c752a0f2f5

Observation 8edfdc6e-430b-43eb-b4c3-34f7e275f8d4 · outbound

This paper cites an unresolved cited work.

Hidden in plain sight: VLMs overlook their visual representations Unresolved cited work

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.971733Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.971733Z digest=sha256:fba712c3c1b4e087d25844631c3073acb864a4e176ce20f9c94d66e240b5c61b

Observation f879e721-808e-4138-b4c2-9b7742b2229a · outbound

This paper cites A, B, C, D.

Hidden in plain sight: VLMs overlook their visual representations A, B, C, D

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-07T05:25:31.975509Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-07T05:25:31.975509Z digest=sha256:6bd4f218426b76e91a017d349fd9a317a7e996c3f72cb64acea5c800ceabe324

Pith citing papers

Observation db4f2ae1-7184-47fd-8352-468f01ff091c · inbound

PictOBI-20k: Unveiling Large Multimodal Models in Visual Decipherment for Pictographic Oracle Bone Characters cites this paper.

PictOBI-20k: Unveiling Large Multimodal Models in Visual Decipherment for Pictographic Oracle Bone Characters Hidden in plain sight: VLMs overlook their visual representations

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-05T05:02:23.812923Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-05T05:02:23.812923Z digest=sha256:99e7edee3f1f978e76a37b5f6a0d9970e04716e2f4e756e2bf1d8fd39eff1125

Observation 3a901a8f-b164-4932-a899-d8b48b76f59b · inbound

SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning cites this paper.

SSL4RL: Revisiting Self-supervised Learning as Intrinsic Reward for Visual-Language Reasoning Hidden in plain sight: VLMs overlook their visual representations

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-21T20:24:21.380035Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T20:24:02.748854Z digest=sha256:ef779d343cb529223f7547f8b13f190515a914b7d86c4c9c74016fb634fe2830

Observation 906c870a-fb4f-47a1-bc78-34f6579d5ef1 · inbound

Mull-Tokens: Modality-Agnostic Latent Thinking cites this paper.

Mull-Tokens: Modality-Agnostic Latent Thinking Hidden in plain sight: VLMs overlook their visual representations

Reference 18

Resolution
verified exact
arxiv_id, observed 2026-05-16T22:58:38.582363Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-16T22:57:04.802871Z digest=sha256:9e1eeba51bad0df717058d3a9bc8b5fa6346cc94167ee6e51f33717c48cfa1c2

Observation 0ece36ae-f997-4c5f-aced-1b408dbee82a · inbound

Egocentric Bias in Vision-Language Models cites this paper.

Egocentric Bias in Vision-Language Models Hidden in plain sight: VLMs overlook their visual representations

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T03:04:42.981300Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T03:04:42.981300Z digest=sha256:4471d23a94c2af3e4a9a564b074ae3644d1274cbc7ad621ce7ae1dc69582f5a6

Observation e02c4924-6ba0-4eec-9be2-0cddcca0ab73 · inbound

Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation cites this paper.

Unifying Contrastive and Generative Objectives for Visual Understanding and Text-to-Image Generation Hidden in plain sight: VLMs overlook their visual representations

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-21T12:20:07.008117Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-21T12:15:57.019720Z digest=sha256:b6ec4da68a29eb2c1ee710aab3b8069143cf749d6402a56656223e2abcca28ba

Observation 843cc588-bcef-45b4-ad02-e53dfa76a3d0 · inbound

Vision Language Models Cannot Reason About Physical Transformation cites this paper.

Vision Language Models Cannot Reason About Physical Transformation Hidden in plain sight: VLMs overlook their visual representations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-15T13:27:51.848177Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-15T13:27:51.848177Z digest=sha256:06ee498a7c85a0d6879acab2bcef77ebc4ca6b76c6b9d44118d787b5558f2efa

Observation 4825b433-659f-4773-8bee-187497140f64 · inbound

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors cites this paper.

VLMs Need Words: Vision Language Models Ignore Visual Detail In Favor of Semantic Anchors Hidden in plain sight: VLMs overlook their visual representations

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-13T21:58:19.843334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-13T21:56:29.924506Z digest=sha256:f46f09a8f62b2a08a574d01ed08af12da2f18d5cba671244fc96dfc55cbce2b7

Observation e41c6df5-c7aa-4fd7-9275-eec59f7b92e8 · inbound

Watch Before You Answer: Learning from Visually Grounded Post-Training cites this paper.

Watch Before You Answer: Learning from Visually Grounded Post-Training Hidden in plain sight: VLMs overlook their visual representations

Reference 22

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T22:20:46.803606Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T20:01:13.305374Z digest=sha256:6d104b531e4f22507efb782636f4b0b5d0936bb8995d526732e3f99a8b0cd21e

Observation deee1951-098a-4113-85f5-857b9badc27d · inbound

Lost in the Hype: Revealing and Dissecting the Performance Degradation of Medical Multimodal Large Language Models in Image Classification cites this paper.

Lost in the Hype: Revealing and Dissecting the Performance Degradation of Medical Multimodal Large Language Models in Image Classification Hidden in plain sight: VLMs overlook their visual representations

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-11T05:21:00.893146Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T18:11:13.523021Z digest=sha256:c307953ef239919ccbaf979296b285af903fddd9ff064000ba4e1b0fbb51240f

Observation 934e2904-d7e1-4c89-978d-3a8e55ac61cf · inbound

Do Vision Language Models Need to Process Image Tokens? cites this paper.

Do Vision Language Models Need to Process Image Tokens? Hidden in plain sight: VLMs overlook their visual representations

Reference 5

Resolution
verified exact
arxiv_id, observed 2026-05-11T00:35:50.692738Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T18:26:37.371488Z digest=sha256:1fa315d399b0b40b66fece748c8e50f5c2cd4cbd2bd4e9f00877a7d13411e505

Observation 2c8aec45-5da4-4685-9d43-3c26b0e35eca · inbound

Boosting Visual Instruction Tuning with Self-Supervised Guidance cites this paper.

Boosting Visual Instruction Tuning with Self-Supervised Guidance Hidden in plain sight: VLMs overlook their visual representations

Reference 25

Resolution
metadata mismatch
arxiv_id, observed 2026-05-11T08:50:58.634642Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T16:27:52.208827Z digest=sha256:2a23504f427310f7ca6d5574534752a1529b32ea5649fc231d21916ce81d4f45

Observation df2215f2-cab2-497c-97a8-f12aad1d624d · inbound

The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models cites this paper.

The Cost of Language: Centroid Erasure Exposes and Exploits Modal Competition in Multimodal Language Models Hidden in plain sight: VLMs overlook their visual representations

Reference 4

Resolution
metadata mismatch
arxiv_id, observed 2026-05-10T13:35:26.811566Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-10T13:25:27.762910Z digest=sha256:1bc16e9d47c0c3338237ed58fda86ca8048c780d533e69396eb1b0addc6d9e00

Observation 575432cf-dadc-494c-807f-1a53dc2b3433 · inbound

Do multimodal models imagine electric sheep? cites this paper.

Do multimodal models imagine electric sheep? Hidden in plain sight: VLMs overlook their visual representations

Reference 24

Resolution
verified exact
arxiv_id, observed 2026-05-12T03:26:19.524703Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-12T03:24:01.933339Z digest=sha256:c32231cd889e76dc63014c8658494ee820c79dc8285096c9218bddcce2b9ccad

Observation 9dc668a8-f105-4526-8b89-61ae92877cf9 · inbound

A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning cites this paper.

A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning Hidden in plain sight: VLMs overlook their visual representations

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-14T19:17:51.063386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-14T19:14:03.809826Z digest=sha256:a19b0fd4298b3dee80795ca84c274af744daa4b114ba1c25331d8ce2d575d468

Observation dc64514a-24f0-456a-b219-b2dbf321a2ba · inbound

A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning cites this paper.

A$_3$B$_2$: Adaptive Asymmetric Adapter for Alleviating Branch Bias in Vision-Language Image Classification with Few-Shot Learning Hidden in plain sight: VLMs overlook their visual representations

Reference 9

Resolution
verified exact
arxiv_id, observed 2026-05-20T22:09:07.103332Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-05-20T22:07:35.021986Z digest=sha256:05c87930f67ec786ec7c40a75ccd6fa928370ce16d9c38c974df735ff8a10c43

Observation b957b92a-4bcf-4ae6-8c5e-30ce662dcb60 · inbound

Diagnosing Visual Ignorance in Vision-Language Models cites this paper.

Diagnosing Visual Ignorance in Vision-Language Models Hidden in plain sight: VLMs overlook their visual representations

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-07-02T16:17:09.193506Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=arxiv_source observed=2026-06-27T22:51:25.269986Z digest=sha256:e43dd908eceb9af4fe5365824559ec239ba4980d3e760756311b327a3086fd05

Observation 43114add-9cbc-41f2-99ff-81fde306c970 · inbound

The Hidden Evolution of Disguised Visual Context inside the VLM cites this paper.

The Hidden Evolution of Disguised Visual Context inside the VLM Hidden in plain sight: VLMs overlook their visual representations

Reference 15

Resolution
metadata mismatch
arxiv_id, observed 2026-07-04T03:29:29.190447Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-06-26T18:08:56.044278Z digest=sha256:2e697c36ba5f73d1064afa567d08c11983693dd2e165663a1704c75a27546ddc

Observation 4ab637ef-2e2c-4309-8082-436d893e5348 · inbound

Visual Access Boundaries in Vision-Language Model Reasoning cites this paper.

Visual Access Boundaries in Vision-Language Model Reasoning Hidden in plain sight: VLMs overlook their visual representations

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-02T06:24:27.325903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:24:27.325903Z digest=sha256:342c79472275cd3a0d934f77f012cfe4af7aa5ca15afe7cfbedb870e0da3d04a

Observation 37f2b892-f65d-4f11-b6cd-7fa3ae1c5d04 · inbound

Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning cites this paper.

Mitigating Visual Degradation in MLLMs via Spatial-Spectral Visual Anchor Learning Hidden in plain sight: VLMs overlook their visual representations

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-15T15:10:13.066770Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T15:10:13.066770Z digest=sha256:1bde6ffadff43bef5143522b5ffb421126cde42bd2e6a3e04a77c5ac2dd18c44

Observation 3fd8021e-4afd-4b44-89b7-8baeb11f729c · inbound

Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO cites this paper.

Teaching MLLMs to Say No: Generalized Referring Expression Comprehension via Refusal Calibrated GRPO Hidden in plain sight: VLMs overlook their visual representations

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T18:44:45.178051Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-08-06T18:44:45.178051Z digest=sha256:04134290a1140dadf032ece3c7d0543164b6225f8944421f95a77c227c45b1c8