Pith. sign in

Paper Citation Record · LEDGER

BLINK: Multimodal Large Language Models Can See but Not Perceive

As of 5 August 2026, this Paper Citation Record lists 90 of 90 outbound references and 60 inbound Pith citation observations for arXiv:2404.12390.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2404.12390 v4

Coverage vector

measured 90 of 90 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-15T20:18:15.439163Z

measured 150 of 150 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-05T06:32:48.257954+00:00

measured 60 of 60 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-03T23:08:49.532620Z

measured 1 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-08-05T02:28:24.338817Z

Reference resolution

90 of 90 outbound references displayed

  • verified exact4
  • verified fuzzy38
  • unresolved18
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch29

External citation measurements

2
pith, observed 2026-08-05T02:28:24.338817Z

Outbound references

Observation fcf71cd7-bd5a-48a6-857f-95277b0fa783 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 1

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.762219Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:6c898513843d82107559d23aa46d8fff205046c37bda67c4253738067028efd4

Observation 93d55819-5eef-47df-922c-cecf99ea015b · outbound

This paper cites In: AAAI (2019) 10.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: AAAI (2019) 10

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.768026Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:c0c17458d85002546456055a06574abecd8129fd889abf7c3d247a45cf286d6c

Observation dfe95755-cdf2-487f-b0d0-d6ae0c098bae · outbound

This paper cites Advances in Neural Information Processing Systems35, 23716–23736 (2022) 2, 4, 22.

BLINK: Multimodal Large Language Models Can See but Not Perceive Advances in Neural Information Processing Systems35, 23716–23736 (2022) 2, 4, 22

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.771138Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:9d5dfc1d7896b2e66a3f44a94c6e752328803155b8fe92cbd0290b1824f13c99

Observation f55ea2a4-f1c8-482a-b894-d35b1898c965 · outbound

This paper cites In: Proceedings of the IEEE international conference on computer vision.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: Proceedings of the IEEE international conference on computer vision

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.773873Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:691dd2f8430a1178c7d897256cee1e472a3644ad9d99216ca849fddee005fd85

Observation d50abe5a-8d1e-41da-a66e-4df0232e8cf2 · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

BLINK: Multimodal Large Language Models Can See but Not Perceive OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.666166Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:0d20c9a9a96e6cc216a5240d2d69909c2185adcaa7aa1d3d362dc1df4b465da6

Observation b698a4e8-b64d-4f81-ae13-8a1661b32373 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 6

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.776403Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:b0ae5e974e88335be60eac9f901a692fb84c0c2455028f7272824bcbf27bae4d

Observation 52168342-21d1-4b73-8c7e-b36acd46bc7c · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

BLINK: Multimodal Large Language Models Can See but Not Perceive Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 7

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.619179Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:c854bb88c1ebf71c698653d74ba7eec48570ee7641bd804ebf508d06a1b37f13

Observation a4b4fa38-fef0-4f95-bd78-9c2c825e419d · outbound

This paper cites In: CVPR (2017) 3, 7.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: CVPR (2017) 3, 7

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.778804Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:9df6e2cd70e62e3964bb68dd7d654418c023ba9f0c29001b71d733f5d76c7c7b

Observation 5a7b5cd3-a7f7-4fac-84c2-11fe68dee4bc · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 9

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.781036Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:647bc57c1194ba3e846b4c95ee11107257b1b0ed3e02bc0ef99b95047afb90d0

Observation 7254d039-d974-4a4c-bc93-e12cd269cf20 · outbound

This paper cites ACM Trans.

BLINK: Multimodal Large Language Models Can See but Not Perceive ACM Trans

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.783620Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:67c4180928a6bd7ce0e6133ced5cb67096147121c0a3f656b91877d2fcb9e59c

Observation b8bad57a-08cb-4d1f-a0c4-6a097a15bc29 · outbound

This paper cites Towards Language Models That Can See: Computer Vision Through the LENS of Natural Language.

BLINK: Multimodal Large Language Models Can See but Not Perceive Towards Language Models That Can See: Computer Vision Through the LENS of Natural Language

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.599633Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:f417ceff084f8b734f32cd0385440184ffc22c0f9819e3fed315686eb79d40a3

Observation b72a520a-b593-4a50-a263-629843d8af92 · outbound

This paper cites In: 1993 (4th) International Conference on Computer Vision.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: 1993 (4th) International Conference on Computer Vision

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.786158Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:081ca05a12474da3e742c93413a870eec981bd993af1f35019f424c8ae55b52c

Observation fed26eb1-3b7d-43be-a90e-39e4331dfbbb · outbound

This paper cites Advances in neural information processing systems33, 1877–1901 (2020) 4, 9, 11, 21, 22.

BLINK: Multimodal Large Language Models Can See but Not Perceive Advances in neural information processing systems33, 1877–1901 (2020) 4, 9, 11, 21, 22

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.788689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:7a2165a9dee6980b064f268a6c3d58c56503525f3340da148b997316d5bf5a35

Observation a9fe8895-9705-4417-a3fb-5a19979de761 · outbound

This paper cites ACM Trans.

BLINK: Multimodal Large Language Models Can See but Not Perceive ACM Trans

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.791003Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:6ef6850b358a7a4316ca25f98705fe796a48cc4fbf8f2c75ce62428d9c573190

Observation 2c66a042-f3bd-4b04-9676-7d13b331b5d2 · outbound

This paper cites In: CVPR (2021) 4.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: CVPR (2021) 4

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.793107Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:e005076720cf89dae929b59a0afd3582767fa1d2d90514d301e66fbc6d327b38

Observation ac9a7f0f-0c40-48eb-83fd-5225b857aa86 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.795278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:01538a5bf3aa3e7d87c6d4e4b971e84067963f73ce8e7675e0e28050696a7960

Observation 996ca3fc-1ad2-4bbe-8cf0-7c41edf64524 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

BLINK: Multimodal Large Language Models Can See but Not Perceive MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 17

Resolution
metadata mismatch
arxiv_id, observed 2026-05-16T07:13:09.112362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:2e8d5aa8debbc73e3c58da8ba369a85ce2061328c0f96bd76424186ea6107a33

Observation 146fed86-e719-4157-a572-2704ba86e6f1 · outbound

This paper cites ShareGPT4V: Improving Large Multi-Modal Models with Better Captions.

BLINK: Multimodal Large Language Models Can See but Not Perceive ShareGPT4V: Improving Large Multi-Modal Models with Better Captions

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.593758Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:37a923b7e888982f2beba1b6d40e997d915e208d948fcc4bae3535aab7b648d0

Observation 46b2e885-baad-45ce-85c9-8704c2eed1ce · outbound

This paper cites Advances in neural information processing systems29 (2016) 3, 7, 8.

BLINK: Multimodal Large Language Models Can See but Not Perceive Advances in neural information processing systems29 (2016) 3, 7, 8

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.798544Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:ea0432d93559e7bb832b2ea913d63767ca434212f835e6746b091118e693ea3f

Observation 8cb85d00-298c-49aa-a09e-3e103837ffbe · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.800932Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:2c0817b41ccdee819bd5aa1f5147d612eafc4e20ec57d77bfd422e45297b0eab

Observation 21afd6bf-0673-4a84-9c75-5b750877a0df · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 21

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.803360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:8c8e765f41998d1727e64d7ddd9c7117e62a8a1dbaa8def1fafbe182aae12823

Observation 998556e5-a872-49a4-9278-fa654d61a263 · outbound

This paper cites https://github.com/open-compass/opencompass (2023) 11.

BLINK: Multimodal Large Language Models Can See but Not Perceive https://github.com/open-compass/opencompass (2023) 11

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.805845Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:63c4aa2012abf69d78fa1726432aadf288d7288accd6d74e8a855e5fae5f4b8e

Observation 0d8d7a4d-c62a-491b-9f6b-a5d314d42318 · outbound

This paper cites com/InternLM/xtuner (2023) 11, 12, 23, 24.

BLINK: Multimodal Large Language Models Can See but Not Perceive com/InternLM/xtuner (2023) 11, 12, 23, 24

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.808247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:6ef64dbb06378fd1cd1618ad4b5f6c886f199ff24f1b48e362c4cb3b34e0835e

Observation 267c7f17-6c3d-449f-9439-fc0dd8361675 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.810690Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:1f2925dec9350647d8d10e1820bc33775ee5cbe846caf953e1f6283d2e7ebb70

Observation b890ceed-b7cc-4a19-999e-3d6fd1dc3926 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 25

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.813663Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:7d85ea62f83c11e5eb2f697b8b5b57311f698cc5fd923da4c9ec3d34d8528071

Observation d4b6a04b-d8b2-46a4-8867-e734fd207a15 · outbound

This paper cites InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model.

BLINK: Multimodal Large Language Models Can See but Not Perceive InternLM-XComposer2: Mastering Free-form Text-Image Composition and Comprehension in Vision-Language Large Model

Reference 26

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T05:30:28.063870Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:67785f916e897ae7adf821ef717d0da60d2e9e89d2a7746ab05a53f97e136b59

Observation c210f13f-37f6-4e5f-abb1-8a52e32644d6 · outbound

This paper cites In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.816068Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:b0c45c462dd3f75cd1a7cfe1bcdace52d106a6feb9035805f1d1439e9c476e2a

Observation 881999c7-fa79-45d0-b917-f348e6369c1d · outbound

This paper cites MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models.

BLINK: Multimodal Large Language Models Can See but Not Perceive MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.509047Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:987d899fb7244f28d6c706bad4f8629c9fd05e93da88d29dc9529632b9d43788

Observation 5001d30e-24ca-480d-828c-8cacd8c79a23 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 29

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.818380Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:d240b2c83e719c29b7402115da809adf18dfc6eb00d9b5235b856524724580f7

Observation 5d08287a-8053-4e97-8d0d-dfe69ce9caae · outbound

This paper cites In: Findings of the Association for Computational Linguistics: ACL 2023.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: Findings of the Association for Computational Linguistics: ACL 2023

Reference 30

Resolution
malformed identifier
doi_truncated, observed 2026-05-15T20:18:15.490957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:611b13f60ab49f580e8006d950ef59aff5283fe3043be1c694214ca27d19a093

Observation 5a80d92d-f532-42ca-83d1-312f60f4cf4d · outbound

This paper cites 37 Wenyue Hua, Lizhou Fan, Lingyao Li, Kai Mei, Jianchao Ji, Yingqiang Ge, Libby Hemphill, and Yongfeng Zhang.

BLINK: Multimodal Large Language Models Can See but Not Perceive 37 Wenyue Hua, Lizhou Fan, Lingyao Li, Kai Mei, Jianchao Ji, Yingqiang Ge, Libby Hemphill, and Yongfeng Zhang

Reference 31

Resolution
metadata mismatch
doi, observed 2026-05-15T20:18:15.485601Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:27147263787ef28d7ab092faf6b480c170919fe5d2da727a941844c72a1ca837

Observation 95663a59-111c-4a1c-9e73-ba265c6dcc19 · outbound

This paper cites Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answering.

BLINK: Multimodal Large Language Models Can See but Not Perceive Dynamic Clue Bottlenecks: Towards Interpretable-by-Design Visual Question Answering

Reference 32

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.582600Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:7ec3309eceecc31e3af117fe72659a92ca241ec03c82c8da9b6b705a9f62ebd6

Observation d7c261a0-478f-475b-b63f-e59ab1eb3003 · outbound

This paper cites In: Conference on Computer Vision and Pattern Recognition (CVPR) (2017) 4.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: Conference on Computer Vision and Pattern Recognition (CVPR) (2017) 4

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.821099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:5fe4c8b2f4113cf1ac651e48a12a6f440be3bd08fbd31f328ab751c631396196

Observation 2d3701d7-a80e-4cd3-a109-1f6dd3636a23 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.823656Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:e2e7a24944283638107947ea8cdcd77e4807678816cde7a79e883b0a46dadd4e

Observation 9cdd2190-f815-4841-9bcf-ced8e62c9fe2 · outbound

This paper cites In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.825953Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:2c1f785d69a602f9cac6bd242777e98b9275cd7f5023559a57d593502b2cdf1b

Observation 0bfe30f8-de75-4de1-a072-3d0165bfd8df · outbound

This paper cites In: Alvey vision conference.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: Alvey vision conference

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.828507Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:7bbe1a760ab40f60ed7021b508550830553334404f1dc3129945f53f619fdba8

Observation a030bb59-3043-41cb-89f5-d0b525733949 · outbound

This paper cites Cambridge university press (2003) 2.

BLINK: Multimodal Large Language Models Can See but Not Perceive Cambridge university press (2003) 2

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.831015Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:52722921b6b44ac20cef0888f6987417e0d46764e291584e9c265b5f73e462c8

Observation 9a16ecd1-667b-4023-b391-602484d52b8b · outbound

This paper cites PromptCap: Prompt-Guided Task-Aware Image Captioning.

BLINK: Multimodal Large Language Models Can See but Not Perceive PromptCap: Prompt-Guided Task-Aware Image Captioning

Reference 38

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.672292Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:34ac7e7232fe52b521d771b940b63ecc914b6c3895fdfb37c78745b2535b871f

Observation 532dd131-098b-434e-a28d-6f80c7074153 · outbound

This paper cites TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering.

BLINK: Multimodal Large Language Models Can See but Not Perceive TIFA: Accurate and Interpretable Text-to-Image Faithfulness Evaluation with Question Answering

Reference 39

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.689615Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:d588f73829a5a2391bf1f13d543e2534da2997edde02133836180bb22e6af0e6

Observation 3b0a7b7e-1c97-493d-b1ad-5a7ae2346f39 · outbound

This paper cites Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models.

BLINK: Multimodal Large Language Models Can See but Not Perceive Visual Program Distillation: Distilling Tools and Programmatic Reasoning into Vision-Language Models

Reference 40

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.496968Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:50b268dd4f295d23dfc2d3c799ce92fde6d22b2f745d1b437479ede7a271f6c7

Observation 0b5a1051-327b-4ebf-8c93-0762dfdced0b · outbound

This paper cites International journal of computer vision123, 32–73 (2017) 3, 4.

BLINK: Multimodal Large Language Models Can See but Not Perceive International journal of computer vision123, 32–73 (2017) 3, 4

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.833614Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:e4204c168fbf46d9a8728cd3d5f48b7ac518eef21d2902a39ef6a4641a2a965c

Observation 7e6c7465-7eff-4c68-ae11-60ebfed00596 · outbound

This paper cites In: Proceedings of the IEEE/CVF International Conference on Computer Vision.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: Proceedings of the IEEE/CVF International Conference on Computer Vision

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.836030Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:8e5e1d2d9419fa787686ffda809ebd17e014f5eb14c3407bb1050288293a655c

Observation 1d5c8533-1aa0-446e-9863-d49c8d67f1cb · outbound

This paper cites SEED-Bench-2: Benchmarking Multimodal Large Language Models.

BLINK: Multimodal Large Language Models Can See but Not Perceive SEED-Bench-2: Benchmarking Multimodal Large Language Models

Reference 43

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.530071Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:53edc9f8bb95e1a4cb776c5504b5a12feea7bdefd6362245c60b0cc98a470fff

Observation 68ea6d96-69d5-4450-8375-3014c5517409 · outbound

This paper cites SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension.

BLINK: Multimodal Large Language Models Can See but Not Perceive SEED-Bench: Benchmarking Multimodal LLMs with Generative Comprehension

Reference 44

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.545662Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:0657aa759ed480b07d229d08f3c35271276da42584b0202c8a1310a790dcf27c

Observation c1412af2-b079-464d-bc52-db91f758aca3 · outbound

This paper cites BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models.

BLINK: Multimodal Large Language Models Can See but Not Perceive BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and Large Language Models

Reference 45

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.558098Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:8db85e4059ce091208e7a693dee57a4086a5d1e06d1dcacbe31c0c3f0425a8aa

Observation a6cadab9-e8d2-42c2-9ed3-29cf7110eab7 · outbound

This paper cites In: Computer Vision– ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: Computer Vision– ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part V 13

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.839038Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:509c91f299956013795c7e998129a8358a05b53b303b541fc15e01b53df16a49

Observation a2d66600-7c1c-4a2f-9eaf-5685180dfec8 · outbound

This paper cites Transactions of the Association for Computational Linguistics11, 635–651 (2023) 2, 9, 21.

BLINK: Multimodal Large Language Models Can See but Not Perceive Transactions of the Association for Computational Linguistics11, 635–651 (2023) 2, 9, 21

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.841877Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:c458a29e9b94ca8912194fb1eab866fcc2d19907a9776a975e91964d2ed328db

Observation 2da219b6-9c93-4953-8dcc-6bc7fbcd1f54 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 48

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.844279Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:210b022ea57839978977daee8f15c1dc18a83e745abbc35c92c21351cc4131da

Observation 14f5665c-0157-4308-98c8-275c83c14c88 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 49

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.846689Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:98a5402f3e48919a0f7cdeac38eb168e9629535f10c1f421dd8d8807ee354946

Observation 509e345e-cd95-4443-b391-bee8bc86f682 · outbound

This paper cites io/blog/2024-01-30-llava-next/ 2, 4, 8, 11, 12, 23, 24.

BLINK: Multimodal Large Language Models Can See but Not Perceive io/blog/2024-01-30-llava-next/ 2, 4, 8, 11, 12, 23, 24

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.849407Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:e152dfbcff6226310ba34dcbc38586406d6279a864cabc2cbd3a7c0f21d30890

Observation 1dd3f07d-6473-444b-97a2-667afbffde6c · outbound

This paper cites Advances in neural information processing systems36 (2024) 2, 11.

BLINK: Multimodal Large Language Models Can See but Not Perceive Advances in neural information processing systems36 (2024) 2, 11

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.693679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:f2db818940427d90251ee46e3e420b9500d86f259dbf5651dad4720b8577122b

Observation cf854146-2f12-47d1-acae-ff5f845c4577 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 52

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.697244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:9a9b2bad288427a7f6b37ce593084bce778aea20bde1bc057f2cf552e755c05f

Observation 84113f97-4b89-46be-8fbb-382d41d8eb83 · outbound

This paper cites OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models.

BLINK: Multimodal Large Language Models Can See but Not Perceive OCRBench: On the Hidden Mystery of OCR in Large Multimodal Models

Reference 53

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T09:55:36.035986Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:80cd8d03019353c3badf8ebd1e01db61bfe4aa74708aaef5d479a89dfd0a78ea

Observation 434afe2a-894c-47b0-ab3b-eaacd2eeb993 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 54

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.700983Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:aa1caf341a8d65226ec02b0e697b3a1a7d170760075c29cc834c0f4327602e02

Observation a66033a4-545f-455c-9f27-947c3b512e4d · outbound

This paper cites In: Proceedings of the seventh IEEE international conference on computer vision.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: Proceedings of the seventh IEEE international conference on computer vision

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.704616Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:c9171ecec4ac9778153afe1c10287eb0d6d548deb40034a99213ee8aabaadc70

Observation 7673fcf4-7502-4ced-8300-5b3a687e1033 · outbound

This paper cites Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unified-IO 2: Scaling Autoregressive Multimodal Models with Vision, Language, Audio, and Action

Reference 56

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.684545Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:f5a95078655bbc57e912e610deda788d22c05d81f12f0001f5ce4b3fd24e17e4

Observation 827638ea-6485-465e-b8cb-f6601dc6fd4e · outbound

This paper cites MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts.

BLINK: Multimodal Large Language Models Can See but Not Perceive MathVista: Evaluating Mathematical Reasoning of Foundation Models in Visual Contexts

Reference 57

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.678753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:d13967fd1d2141bc68e790317988e24ced731674cbe0d4b0b4205859c0021ba1

Observation cc83e521-21b9-41d8-9c3e-aa999faf7126 · outbound

This paper cites MIT press (2010) 2.

BLINK: Multimodal Large Language Models Can See but Not Perceive MIT press (2010) 2

Reference 58

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.709099Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:c58d8fc349193eead020cef84c62441e3afc6b2c3426652cfc644d2c65f45560

Observation 0eeac0e1-ac71-45ef-9a00-80fac5c18bb4 · outbound

This paper cites Science 194(4262), 283–287 (1976) 2.

BLINK: Multimodal Large Language Models Can See but Not Perceive Science 194(4262), 283–287 (1976) 2

Reference 59

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.713373Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:12f940b4f5b07350905c592c0330201a4c529b9e048fc7968a277c88e001a113

Observation f49f3f48-d0ac-4389-8315-31b5ccd71c56 · outbound

This paper cites SPair-71k: A Large-scale Benchmark for Semantic Correspondence.

BLINK: Multimodal Large Language Models Can See but Not Perceive SPair-71k: A Large-scale Benchmark for Semantic Correspondence

Reference 60

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.503362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:da960f2e5f1d7eb281f83da07bece4574f866cb7a978b7eb12dc17cd11f60188

Observation bfb6c897-1a2e-4c19-be5e-fd2aee2e9fd9 · outbound

This paper cites Cambridge tiass., HIT479(480), 104 (1969) 2.

BLINK: Multimodal Large Language Models Can See but Not Perceive Cambridge tiass., HIT479(480), 104 (1969) 2

Reference 61

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.717033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:5d3357b742c54992da1e35b135fddb342f420b55070caf4fe4117efd1be84883

Observation 1ed26a87-e17c-4207-8559-e3d004ac42b5 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 62

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.720719Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:b393de432f7252be99528f10d40fecb6139715680c16da5afec993d1827caeba

Observation f91b33e7-a376-4ba9-ac0e-49845d32d7ce · outbound

This paper cites SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis.

BLINK: Multimodal Large Language Models Can See but Not Perceive SDXL: Improving Latent Diffusion Models for High-Resolution Image Synthesis

Reference 63

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.522936Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:7b33365565161ffa65343b3cdb80e3eb4519d448f22dc8fb99edc5dbf75045ae

Observation 7ff96374-5266-4210-a5a0-706d7dcf6277 · outbound

This paper cites In: International conference on machine learning.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: International conference on machine learning

Reference 64

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.724244Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:88b2d68b29ab8ab9cb400a32104454c20adc2d4499f4835b4f85b1eb552ea549

Observation 410795db-2cbf-4622-a1a2-34c8ce65d5ff · outbound

This paper cites LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs.

BLINK: Multimodal Large Language Models Can See but Not Perceive LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs

Reference 65

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.537757Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:f927aa15e4cfb67f909543728061ddab5f56071dc4f84d920c6646e272946c95

Observation 33c1b918-6657-4f1a-8e24-48fcad2e8f2f · outbound

This paper cites In: European Conference on Computer Vision.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: European Conference on Computer Vision

Reference 66

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.728266Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:4d29735253649579e91f2019a79ede39a51c2eafcc9a80662d8b4d58828fc041

Observation dcb66970-d321-4245-98b2-a09e0c80b1e4 · outbound

This paper cites What does CLIP know about a red circle? Visual prompt engineering for VLMs.

BLINK: Multimodal Large Language Models Can See but Not Perceive What does CLIP know about a red circle? Visual prompt engineering for VLMs

Reference 67

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.551709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:b9a4e0495e2ab4e7751dcaeee6c463c854472d71081435e897de396daaeb58ce

Observation 5eb3dfad-7ab7-465a-ac49-c3ca3ec61d1f · outbound

This paper cites CVPR (2021) 14.

BLINK: Multimodal Large Language Models Can See but Not Perceive CVPR (2021) 14

Reference 68

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.731452Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:e8668dff49b8356ac3327f97c7c0350ce238b23e22a3fcc10f3d96bf70eb3459

Observation b8d17ed1-6576-4d50-ad15-e05a10d25bd1 · outbound

This paper cites EVA-CLIP: Improved Training Techniques for CLIP at Scale.

BLINK: Multimodal Large Language Models Can See but Not Perceive EVA-CLIP: Improved Training Techniques for CLIP at Scale

Reference 69

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.564212Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:440afdcfd884fce3feb02ae72c8dc09c53bf8ecb1233d33684c793923413a5c0

Observation b4aef595-17b8-4569-bcd5-562b37ff4ffe · outbound

This paper cites Emergent Correspondence from Image Diffusion.

BLINK: Multimodal Large Language Models Can See but Not Perceive Emergent Correspondence from Image Diffusion

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.570881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:7db9a511e7a3122631a5b961b5085f642fd7382e846f42aad6dc4d34d7478899

Observation c17df47c-94db-4fdc-82ff-2839578afde7 · outbound

This paper cites Gemini: A Family of Highly Capable Multimodal Models.

BLINK: Multimodal Large Language Models Can See but Not Perceive Gemini: A Family of Highly Capable Multimodal Models

Reference 71

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.576419Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:1f9075ca400f619a5127cb37c9c2c07f1fe195da9b652f5388fdf69ed684e9bc

Observation a7075c6b-8591-482f-83b5-e0ec6c073911 · outbound

This paper cites https://github.com/InternLM/InternLM (2023) 12, 23, 24.

BLINK: Multimodal Large Language Models Can See but Not Perceive https://github.com/InternLM/InternLM (2023) 12, 23, 24

Reference 72

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.734581Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:87ddeb1213fcb4baf8e92a2a186b5e8f65a455c20166a52589687fd8c8d21f5f

Observation f4cae9cd-991d-4d5b-901a-c596728f644a · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 73

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.737730Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:400633e99e7a68b1a12cac0624d9de22e92f5297eed44f59f475f1c88566fa2d

Observation 2a3aae24-6ba9-4f12-96a2-661cbac603c4 · outbound

This paper cites IEEE Transactions on pattern analysis and machine intelligence24(9), 1226–1238 (2002) 2 20 Fu et al.

BLINK: Multimodal Large Language Models Can See but Not Perceive IEEE Transactions on pattern analysis and machine intelligence24(9), 1226–1238 (2002) 2 20 Fu et al

Reference 74

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.740549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:15b03ad814fc7f9c6c88955819daa338fd5be242da6a9714f29809ca566d81d4

Observation c06fee5a-2d65-4152-a868-6783f875cfff · outbound

This paper cites Llama 2: Open Foundation and Fine-Tuned Chat Models.

BLINK: Multimodal Large Language Models Can See but Not Perceive Llama 2: Open Foundation and Fine-Tuned Chat Models

Reference 75

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.588230Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:2d2f5adfca2ea79b8d78215cb1527b522ae9dc9232a260876cb8b4d56b382e69

Observation 56a1c8a3-f519-442d-a6bb-eb1cd0f04b5a · outbound

This paper cites In: Pro- ceedings of IEEE Conference on Computer Vision and Pattern Recognition.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: Pro- ceedings of IEEE Conference on Computer Vision and Pattern Recognition

Reference 76

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.743471Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:60649b909d17a391b2133c9960795ad3b507ac111d8ce7aa742052f54965eff7

Observation 96aa267c-ca6c-47f5-8249-15ff9908c10e · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 77

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.745998Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:68b7b12dcabea17499fd64f5fdc691656a2d0534fb5a8c827228ed8b181de538

Observation 21050174-c0c5-400c-b14a-73891c3f4531 · outbound

This paper cites Self-Consistency Improves Chain of Thought Reasoning in Language Models.

BLINK: Multimodal Large Language Models Can See but Not Perceive Self-Consistency Improves Chain of Thought Reasoning in Language Models

Reference 78

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.606264Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:ab09dc9cea860825ac9cf4ee94b03385690ba2325d50f84ce42a0fb63868197b

Observation 8a04abfe-161d-49f3-96c3-2728f4fad943 · outbound

This paper cites DIRE for Diffusion-Generated Image Detection.

BLINK: Multimodal Large Language Models Can See but Not Perceive DIRE for Diffusion-Generated Image Detection

Reference 79

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.612995Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:c702930d62f372c574625710d09b2cc7118cd50ae43b2af6c05eb2a7e6231346

Observation fc6a34db-744a-4677-a780-d0a488b920e3 · outbound

This paper cites an unresolved cited work.

BLINK: Multimodal Large Language Models Can See but Not Perceive Unresolved cited work

Reference 80

Resolution
unresolved
raw_fallback, observed 2026-05-15T20:18:15.748558Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:9bd2c50994c231d376fb4f693e54c8c34eda966f69645813aaa7089dfbe14bad

Observation 3be8907e-eadd-4f79-ab9f-c308f9a3827c · outbound

This paper cites List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs.

BLINK: Multimodal Large Language Models Can See but Not Perceive List Items One by One: A New Data Source and Learning Paradigm for Multimodal LLMs

Reference 81

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.625753Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:9b2f61833bcd01bd9bb511abe0da4647d1ab524e42ea124789d0928c78e86028

Observation f3fee5fc-46c8-4076-b40f-91e0f1a6489a · outbound

This paper cites Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V.

BLINK: Multimodal Large Language Models Can See but Not Perceive Set-of-Mark Prompting Unleashes Extraordinary Visual Grounding in GPT-4V

Reference 82

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.632677Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:6f0c50eb69d213e9f20dff82428ebba8e9fc127fd4315128af587814ee21d837

Observation 7fced311-c5dd-45cc-9e03-b3a3fb4f3f82 · outbound

This paper cites In: CVPR (2024) 14.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: CVPR (2024) 14

Reference 83

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.751646Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:0a7ac410f1e33f28362c8f5269c02987ecb7e6da5cf22c93bc3bc57464566b1e

Observation e8d55ce6-fd0d-41eb-80e4-30352dbd407a · outbound

This paper cites In: Proceedings of the AAAI Conference on Artificial Intelligence.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: Proceedings of the AAAI Conference on Artificial Intelligence

Reference 84

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.754381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:fdc457cf4acb046cccfe9c633384f847d274d59f63e5fc68431e664297ffb741

Observation 84fa579c-1abd-48d2-b560-b96da5f3a348 · outbound

This paper cites The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision).

BLINK: Multimodal Large Language Models Can See but Not Perceive The Dawn of LMMs: Preliminary Explorations with GPT-4V(ision)

Reference 85

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T23:26:07.065720Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:68cd5434e44f504b8a0f0b1a57ca7e59b6a8916fbca3054f09096f62f246bb89

Observation 8daddfca-2328-4e99-867c-dfe394bb0e23 · outbound

This paper cites MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities.

BLINK: Multimodal Large Language Models Can See but Not Perceive MM-Vet: Evaluating Large Multimodal Models for Integrated Capabilities

Reference 86

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.655809Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:eb8ef1e6b3f491fe0a7185a00d323d245e9b04400438d06aa85a41e5d7b39b9a

Observation 566ba0cc-e3c3-4f53-9bf6-6bb74f6bc283 · outbound

This paper cites MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI.

BLINK: Multimodal Large Language Models Can See but Not Perceive MMMU: A Massive Multi-discipline Multimodal Understanding and Reasoning Benchmark for Expert AGI

Reference 87

Resolution
metadata mismatch
local_arxiv, observed 2026-05-15T20:18:15.661112Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:f1c3d6771b2a75f919d77d12b1bf7d33c4a41f59ef06988574af7d2e6bbd7f83

Observation d8059fec-5cab-4dab-b8b9-1830dc7cbd2b · outbound

This paper cites Advances in Neural Information Processing Systems35, 27469–27483 (2022) 3, 9.

BLINK: Multimodal Large Language Models Can See but Not Perceive Advances in Neural Information Processing Systems35, 27469–27483 (2022) 3, 9

Reference 88

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.757131Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:6792a93c2eba0784bfcfd687b60c3d595495a7ce4dada5758745c65535cc60af

Observation fa975452-116c-4d8e-860a-9083bc9de309 · outbound

This paper cites In: The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (June 2019) 4.

BLINK: Multimodal Large Language Models Can See but Not Perceive In: The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (June 2019) 4

Reference 89

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.759785Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:7973770f7c1a13f0023610ce68bef76db724d2a8cabc7204732e6a33a3ba7556

Observation f777c19d-b5ac-4ae3-8fec-a230e670e8f9 · outbound

This paper cites You are an AI assistant who will help me to match an answer with several options of a single-choice question.

BLINK: Multimodal Large Language Models Can See but Not Perceive You are an AI assistant who will help me to match an answer with several options of a single-choice question

Reference 90

Resolution
verified fuzzy
raw_fallback, observed 2026-05-15T20:18:15.765110Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-15T20:18:15.439163Z digest=sha256:cd737dad619145d2f8da091841d509b9b12fab1cfa6999bf3c74b34c6616c3b3

Pith citing papers

Observation 449cc6a2-7af6-44e2-8b8c-60d35ee072a1 · inbound

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone cites this paper.

Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T20:19:27.255515Z digest=sha256:aff39fdd98127ed02dd2665b58c627e0c4d79c9ff1955062e22d959697485665

Observation db3428af-e77a-4f8a-995c-184368d595ba · inbound

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding cites this paper.

MuirBench: A Comprehensive Benchmark for Robust Multi-image Understanding BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 17

Resolution
verified exact
local_arxiv, observed 2026-05-17T01:09:30.442631Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T01:09:30.360275Z digest=sha256:7299229f7dc727c5cbc476900a76d14793006f9a66263074d43646391d5db793

Observation f2da540d-4c8c-4aa0-8482-62bf96050759 · inbound

Depth Anything V2 cites this paper.

Depth Anything V2 BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 21

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T14:56:33.945280Z digest=sha256:aceacd13073307cdf19aa5e20ab36fbf75b1820d8b77f28e770ecb1d5c8c69f3

Observation 35fbd366-aacc-42ef-8c5b-3de03abd4298 · inbound

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs cites this paper.

Cambrian-1: A Fully Open, Vision-Centric Exploration of Multimodal LLMs BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 42

Resolution
verified exact
local_arxiv, observed 2026-05-17T00:05:03.712019Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T00:05:03.547664Z digest=sha256:16092dfd86737a109d181f098d9e682cccd38804c5a5f53fdb8542cb8ed23d6d

Observation 2490bbe3-8e77-4a4e-ae5a-95e5937f0c81 · inbound

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models cites this paper.

LLaVA-NeXT-Interleave: Tackling Multi-image, Video, and 3D in Large Multimodal Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T06:01:53.730356Z digest=sha256:17ef2991b148ef8a8890f86e2ec3670ab8fb129f146783880287917d40af2d2f

Observation d256730f-a2e3-4983-90d8-0043a8ba6f24 · inbound

LLaVA-OneVision: Easy Visual Task Transfer cites this paper.

LLaVA-OneVision: Easy Visual Task Transfer BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 31

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T14:23:49.412830Z digest=sha256:dabc7eb39168cbe8565c339de87931bbcdddc06b5913188825bfff228ab5bedc

Observation a910a12c-9f03-4125-b335-711b47e51e93 · inbound

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? cites this paper.

MME-RealWorld: Could Your Multimodal LLM Challenge High-Resolution Real-World Scenarios that are Difficult for Humans? BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 20

Resolution
verified exact
local_arxiv, observed 2026-05-16T07:59:32.699685Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T07:59:32.638758Z digest=sha256:2dc28d9e2609c86100439ba6c5f98999320b15b86db224d2500bee091732f8d3

Observation 02f44911-d49a-470d-b4a6-7515bb7319a4 · inbound

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark cites this paper.

MMMU-Pro: A More Robust Multi-discipline Multimodal Understanding Benchmark BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-14T00:51:48.163349Z digest=sha256:126423353f5477d84b50743c4875db8090b8ba2192d42ea0cc999a06411aaa13

Observation b97e81f7-0bde-406c-b183-6b7e09787a8c · inbound

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling cites this paper.

Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 70

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T13:23:57.588851Z digest=sha256:305dd3177df0480dee4b4af3eb3fc1bb53c84c6ebba325fe31ce2d340149bf95

Observation d9b79575-df07-451c-ab2a-4b28228f0063 · inbound

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning cites this paper.

MetaMorph: Multimodal Understanding and Generation via Instruction Tuning BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 258

Resolution
metadata mismatch
local_arxiv, observed 2026-05-17T07:51:13.207957Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-17T07:51:12.953777Z digest=sha256:1b56830cb59c4837e9bbe3b669f668d4beb7024afd24c65c722a515d12cea507

Observation cb91d072-e5e7-429c-b9d4-f256aecc44af · inbound

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs cites this paper.

Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T22:22:27.455361Z digest=sha256:86311a4645cb298469649c397f922e784bd245e5ca9cd9ed77afa95211a34698

Observation 2637342f-824d-4ee1-8cec-4d9dbe651456 · inbound

Gemma 3 Technical Report cites this paper.

Gemma 3 Technical Report BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 18

Resolution
verified exact
local_arxiv, observed 2026-05-22T22:22:12.223017Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T22:18:55.976503Z digest=sha256:58fd588c182491b621bb5f386d66e56c9ccafba63af39a2e543711283c9523ee

Observation 2a242639-a304-471e-8c5f-5b5ad78f44e1 · inbound

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models cites this paper.

InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 39

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T13:41:07.991012Z digest=sha256:7c3c074b1ee15e21cfbc9fadce506983988c62d512bc0cb5e7994942744aaa27

Observation 3024c293-8945-492a-9bb2-c065a21ddb9c · inbound

Grounded Reinforcement Learning for Visual Reasoning cites this paper.

Grounded Reinforcement Learning for Visual Reasoning BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-05-22T01:05:52.174278Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-22T01:05:18.801388Z digest=sha256:68fe9b54975a675cb26539d0433a9b1db0804639a43f70dfad366c733b2df342

Observation 7095c77d-5117-459b-ac10-881acf5f2746 · inbound

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning cites this paper.

GLM-4.5V and GLM-4.1V-Thinking: Towards Versatile Multimodal Reasoning with Scalable Reinforcement Learning BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-11T04:48:26.355351Z digest=sha256:57519a05a9f2bd3c7752e266fd67cd3f46b5c0f7f01606cfe81c8cd89359289e

Observation 84230edf-4209-4308-9141-436bcd19977b · inbound

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks cites this paper.

How Well Does GPT-4o Understand Vision? Evaluating Multimodal Foundation Models on Standard Computer Vision Tasks BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 21

Resolution
verified exact
local_arxiv, observed 2026-05-19T05:57:08.011869Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-19T05:55:09.188048Z digest=sha256:8a66ff98b623ba787f79aa8552017a7cb57fafa3680dc652dfcbece14961146e

Observation d517cdae-32fd-4521-a495-ce6109d7355a · inbound

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency cites this paper.

InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 36

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T11:58:58.660564Z digest=sha256:c7833a80e35fa80b6818e2be2de232ea83c281f5932f69826dfb861cb12a3643

Observation 80da89b1-98fb-43e9-8400-1d37f0472ea6 · inbound

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards cites this paper.

SpatialThinker: Reinforcing Scene Graph-Grounded Spatial Reasoning via Dense Rewards BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T23:08:49.532620Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T23:08:49.532620Z digest=sha256:dd856eba975e368ed7d11d917165f4e0e63523ca6b9b9221c4d43d1faef71b3e

Observation f0d3fd3c-a303-42f7-b729-f26b80345f4b · inbound

Abstract 3D Perception for Spatial Intelligence in Vision-Language Models cites this paper.

Abstract 3D Perception for Spatial Intelligence in Vision-Language Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 11

Resolution
verified exact
local_arxiv, observed 2026-05-17T22:45:24.419680Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-17T22:43:16.761970Z digest=sha256:1c93226c908ee1ccf048150adf2573f2d515278aa926d615b7eb028cb1eefbac

Observation c8e4af68-f378-4907-8e8f-4673229054ca · inbound

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL cites this paper.

SpaceTools: Tool-Augmented Spatial Reasoning via Double Interactive RL BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T18:42:11.109093Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T18:42:11.109093Z digest=sha256:1715b4db6d3e0b39e8bd174d6b41e64fca03c0504af79a03924568ff3e50e9ea

Observation ef67d5fd-d15f-4fda-b4b0-c5ab2d75330a · inbound

Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning cites this paper.

Omni-R1: Towards the Unified Generative Paradigm for Multimodal Reasoning BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 31

Resolution
metadata mismatch
local_arxiv, observed 2026-05-16T14:37:59.940145Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-16T14:37:05.402850Z digest=sha256:282e1205dd5ca000b0be8f612ff50c2af89bbdb7d910531fadd0ea9080763be3

Observation ac112fd3-56c9-4fba-b4f9-128456022928 · inbound

Kimi K2.5: Visual Agentic Intelligence cites this paper.

Kimi K2.5: Visual Agentic Intelligence BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T16:09:05.225767Z digest=sha256:8f494576d9188c57e7d5cc57efa3d54cf33cac63ca06641a46a1eb2a4d22c992

Observation fcab69be-4460-4126-8fe5-a1838387d2b4 · inbound

Multimodal Language Models Cannot Spot Spatial Inconsistencies cites this paper.

Multimodal Language Models Cannot Spot Spatial Inconsistencies BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-13T23:12:27.333405Z digest=sha256:51a65d8cf3ba950840fd93832b0aa57cbc5ba77e0a352cdcd9578814abee3b23

Observation cb7da1c8-5250-4d3a-88a7-3395600d6638 · inbound

EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training cites this paper.

EmbodiedMidtrain: Bridging the Gap between Vision-Language Models and Vision-Language-Action Models via Mid-training BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 7

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-10T02:16:08.687340Z digest=sha256:43da758ae585c860e71729db010743f877724ffcb58e27e35c941c9606d59821

Observation 5c7952db-71b9-4556-b907-5c59a7f6b83a · inbound

Improving Vision-language Models with Perception-centric Process Reward Models cites this paper.

Improving Vision-language Models with Perception-centric Process Reward Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-08T04:33:36.634359Z digest=sha256:b38529b49b3900872bedf2b23ad84b6cbe52a223453bd3921d3095b75d7b5ddb

Observation ed1cadea-6207-4b5c-8aa3-40ecee404002 · inbound

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs cites this paper.

Visual Latents Know More Than They Say: Unsilencing Latent Reasoning in MLLMs BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-09T15:47:49.982564Z digest=sha256:a63900a29e727eb528068c6c24eb4ccfbcc7d799bb0261520aa08413bc410b66

Observation ac8a1ce3-7d20-450c-90f2-fe7fab2e5efc · inbound

RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction cites this paper.

RetentiveKV: State-Space Memory for Uncertainty-Aware Multimodal KV Cache Eviction BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 6

Resolution
metadata mismatch
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-05-10T15:29:17.567557Z digest=sha256:d005c79d964ad559f2012ac8a4b619c9bfe90312d69f6de28a0d27a2fd802c5f

Observation d57cdf37-8b01-47fa-bdff-def804182c4c · inbound

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space cites this paper.

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 13

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-12T04:30:54.053958Z digest=sha256:3179327fc1554f6076924dcc4e157209d7126aeca3d6727e6e19a07ce9087251

Observation 5a4ce0a0-ebc7-45a1-9d80-6b86ae9c5d7e · inbound

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space cites this paper.

The Cartesian Shortcut: Re-evaluate Vision Reasoning in Polar Coordinate Space BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 13

Resolution
verified exact
local_arxiv, observed 2026-07-01T13:35:46.604914Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T22:58:04.574536Z digest=sha256:6fb2b16968ffa4561404e7f41096752acf25a17d81f40de64c3a2b1bc61322ba

Observation 653ca010-be3a-4fd0-8c2c-ccb7b67d391b · inbound

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone cites this paper.

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-13T02:52:43.674969Z digest=sha256:5f5b35bbe252d9d9ffbeadc26420a30c7d18535f0eef187723454928c6a834a0

Observation 5de67ce0-c44b-4ec6-8abc-43c140d464b1 · inbound

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone cites this paper.

20/20 Vision Language Models: A Prescription for Better VLMs through Data Curation Alone BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 15

Resolution
verified exact
arxiv_id, observed 2026-05-15T20:18:15.850579Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-14T21:28:37.680681Z digest=sha256:3ce18c46cba108cdcbcbfa8002712ab58a521a92e0983cde99744d2e51bb2c82

Observation 0680746e-0946-49b5-982f-d46c5df6dcb0 · inbound

When Vision Speaks for Sound cites this paper.

When Vision Speaks for Sound BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 18

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T22:13:46.799961Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T22:12:52.160596Z digest=sha256:19a3e897e4f129af88c4547f983a442516a873316aa8794afd18d54f51c7e591

Observation 32f15706-6d81-4527-a732-c653269f5fe0 · inbound

What's Holding Back Latent Visual Reasoning? cites this paper.

What's Holding Back Latent Visual Reasoning? BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 6

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T12:03:15.479718Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T11:59:14.134917Z digest=sha256:85fde6aa881ce5254614704cd06140a7d7e265c1e34af9e0a337587d271b99d5

Observation 45972e6c-4f3a-49ce-ad25-59e5e90a968d · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-05-20T11:33:14.185570Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T11:32:24.007847Z digest=sha256:5af58ed3050c8f7ab307131a607122b243c81508ae2c0fcd6ab6b7fdd166d67f

Observation 41f4ffe9-b6b8-4ff8-819f-eb9b5000c5c0 · inbound

Semantic Generative Tuning for Unified Multimodal Models cites this paper.

Semantic Generative Tuning for Unified Multimodal Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:35:00.371150Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T18:31:10.578558Z digest=sha256:75b13dc2979e03da3bd51561410e99795324a5272c154ce3b113fbaa93af8586

Observation 7ab9d2bd-131f-4b57-bd9f-16c354b70c5e · inbound

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop cites this paper.

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-05-20T10:53:13.208797Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-20T10:52:22.778489Z digest=sha256:f42590699b5d234557b80654f6c91508dbf753bca84a6163996a2777e4f2e006

Observation c2a322ea-d23f-496d-8dcc-18c43c501de5 · inbound

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop cites this paper.

ESI-Bench: Towards Embodied Spatial Intelligence that Closes the Perception-Action Loop BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 7

Resolution
verified exact
local_arxiv, observed 2026-07-01T15:05:47.176193Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T18:25:17.831116Z digest=sha256:72b9bcb2db41915d7f719c16c8f443e9f5a22599bd5618933cf04d865eba9b09

Observation 7ad6f110-2913-4142-bc38-ab555372727b · inbound

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison cites this paper.

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-05-21T08:39:53.709641Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-05-21T08:36:27.676888Z digest=sha256:8e89385d9c284b41e40e39764d99f9c19d2634a948ee1cc537de859f1414c701

Observation b280e90a-091a-4308-a785-3ce2c8c263d9 · inbound

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison cites this paper.

ClaimDiff-RL: Fine-Grained Caption Reinforcement Learning through Visual Claim Comparison BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 8

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T18:04:58.341972Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T17:57:47.409741Z digest=sha256:67a4c384d29f7d809afaf11e5fffef31f6f73718ad7769d38e63ea809ea560d4

Observation 870f1a3d-4d80-4c90-9713-611b8411423a · inbound

PInVerify: An Offline Embodied Benchmark for Active Instance Verification cites this paper.

PInVerify: An Offline Embodied Benchmark for Active Instance Verification BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 16

Resolution
metadata mismatch
local_arxiv, observed 2026-06-29T07:33:13.552711Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T07:30:59.306097Z digest=sha256:57528e9be611b3d05d9de469f6b710a047a5210658f44b44115b77bc146c4c49

Observation 508d172c-b047-4f6a-9dad-28b2652407c4 · inbound

VisualFLIP: Do Predictions Depend on Task-Critical Visual Evidence in Multimodal Reasoning? cites this paper.

VisualFLIP: Do Predictions Depend on Task-Critical Visual Evidence in Multimodal Reasoning? BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 17

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T17:37:14.727108Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T21:56:25.350827Z digest=sha256:4681851344889b36cfb4101afaf309f4c6b18103cf1047b3c16f5c77c6ae5d37

Observation b148aae3-f0e4-4bae-aa81-fc50c8314a12 · inbound

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models cites this paper.

Embodied-R1.5: Evolving Physical Intelligence via Embodied Foundation Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 31

Resolution
unresolved
no resolver link, observed 2026-07-14T18:07:09.018997Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T18:07:09.018997Z digest=sha256:9db8eba6c0b42bec25ef0a06d657d2ebcd34a21bb47d7884ae5cec723ab7c215

Observation 6351c0a6-442f-44bf-bf28-b82b86efef77 · inbound

Human-Enhanced Loop Modeling (HELM): Agent-Based Finite Element Modeling of Concrete Bridge Barriers cites this paper.

Human-Enhanced Loop Modeling (HELM): Agent-Based Finite Element Modeling of Concrete Bridge Barriers BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 3

Resolution
verified exact
local_arxiv, observed 2026-06-27T09:50:49.101338Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-11T11:50:26.030339Z digest=sha256:76e22f7bf6ce09d8242797e5fc4e5f9c772d3930b76f47218468702340945ce6

Observation 10e1f930-285b-421e-babd-a0e88f519f27 · inbound

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models cites this paper.

Reinforcing Dual-Path Reasoning in Spatial Vision Language Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 126

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T20:08:55.668706Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-27T01:42:30.005911Z digest=sha256:eb64faf51c325e6e8beb27691bdde67bd038946a3b09e9fbde2d3abedd2b9a21

Observation fdc46151-68b7-4d87-a1d1-123208f6c79c · inbound

Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning cites this paper.

Visual-OPSD: Cross-Modal On-Policy Self-Distillation for Efficient Unified Multimodal Reasoning BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 25

Resolution
metadata mismatch
local_arxiv, observed 2026-07-03T23:49:02.264842Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=arxiv_source observed=2026-06-26T21:47:51.437284Z digest=sha256:5198ee2619ebadf66125e4ba0650652cc7f08856db9860d09f04fb56167a7f79

Observation 13ac3131-27ae-4e06-b3cf-71889a166f29 · inbound

S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence cites this paper.

S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T03:29:31.569062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T17:56:07.864580Z digest=sha256:8909510dcb54753af3798380102bb04a92412df519e705af3f02b82fa5dce15e

Observation d41550c2-7965-45d5-9bed-ea092705becc · inbound

S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence cites this paper.

S-Agent: Spatial Tool-Use Elicits Reasoning for Spatial Intelligence BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 11

Resolution
metadata mismatch
local_arxiv, observed 2026-06-30T11:54:39.058350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-30T10:22:08.535453Z digest=sha256:7df581043f988592287caa15b43931dc2be93c9d0304a74de1ee6fb6650e52fd

Observation a7bf10db-3965-469f-9212-8b0376aa8f1e · inbound

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception cites this paper.

One Image is All You Need: Agentic One-Shot Image Generation via Text-Based World Models for Long-Tail Spatial Perception BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 15

Resolution
verified exact
local_arxiv, observed 2026-07-04T03:09:30.076334Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T18:22:41.189820Z digest=sha256:0332c282452ab60bf3ebc8419083345f3ef590edbe32780d87553070f0beffec

Observation 9b6bcb5c-80c8-492c-b2eb-08e16619da25 · inbound

When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR cites this paper.

When Does a Video-Language Model Stop Watching? Reward Strength Controls the Formation and Reversal of Visual Shortcuts in Multimodal RLVR BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 2

Resolution
verified exact
local_arxiv, observed 2026-07-04T08:29:41.287801Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T11:43:17.464276Z digest=sha256:6bd01d971c57bd3754f344713810307b4c00880dd4a7e47e88c13fa21f37b568

Observation 65ca5b1c-22f4-4574-a1e6-af195f0604d6 · inbound

TuringViT: Making SOTA Vision Transformers Accessible to All cites this paper.

TuringViT: Making SOTA Vision Transformers Accessible to All BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 58

Resolution
verified exact
local_arxiv, observed 2026-06-29T15:03:32.200965Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-29T05:32:26.746776Z digest=sha256:60a11325def8325305b592fa1ca7489ed3319ce5f3111a12c87583390fe1cc60

Observation f0acdfa3-2518-438c-9d35-bf02dd304185 · inbound

Do vision-language models search like humans? Reasoning tokens as a reaction-time analog in classic visual-search paradigms cites this paper.

Do vision-language models search like humans? Reasoning tokens as a reaction-time analog in classic visual-search paradigms BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 5

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T18:10:02.344679Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-25T22:58:41.991573Z digest=sha256:b49240f3ce18334e1cac47cb2bd3ebdd58b246cf7dae8e9d911c4e45a1d55fba

Observation f79a77b4-41dc-4519-8def-0f5e1e5622ac · inbound

C3-Bench: A Context-Aware Change Captioning Benchmark cites this paper.

C3-Bench: A Context-Aware Change Captioning Benchmark BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 28

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T19:50:10.197607Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-25T21:02:52.529391Z digest=sha256:373b8f58e29a42e67ad8588b1314ea54e20f835a0ac6c76e86fde0e862d2266f

Observation 5ae8d44e-6985-4875-b221-1261d8c47ec4 · inbound

DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues cites this paper.

DiCoBench: Benchmarking Multi-Image Fine-Grained Perception via Differential and Commonality Visual Cues BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 9

Resolution
metadata mismatch
local_arxiv, observed 2026-07-04T13:19:50.970106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-06-26T05:18:01.931929Z digest=sha256:362bf6767e2e451dba85570e615656b805b402a53c6fafa0ea113757e3aca666

Observation 6265e2a3-e2ac-48c7-afb9-a4c84b86d857 · inbound

Safe Alone, Unsafe Together: Safeguarding Against Implicit Toxicity When Benign Images Combine cites this paper.

Safe Alone, Unsafe Together: Safeguarding Against Implicit Toxicity When Benign Images Combine BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 1

Resolution
metadata mismatch
local_arxiv, observed 2026-07-02T13:26:58.578997Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-05T06:32:48.257954+00:00.

source=pdf_text observed=2026-07-02T13:19:29.775959Z digest=sha256:34710a550aee857723a3c18745f0c5d856aef58430cfcc5c390f03b1fbb2caac

Observation ae5cf0f9-eb9a-47b2-9552-76e22787a979 · inbound

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space cites this paper.

ProLaViT: Learning Progressive Latent Visual Thoughts in Structured Latent Space BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T06:13:16.894125Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T06:13:16.894125Z digest=sha256:5852e864554b9795d388a1a638e304e663be762ce2cf40b76f738a8e08acb99c

Observation 21908eb4-c7c1-4ead-a2f4-43266c6052d7 · inbound

SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning cites this paper.

SVR-R1: Bootstrapping Multi-modal Reasoning with Self-verification in Reinforcement Learning BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-14T07:59:56.441098Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=arxiv_source observed=2026-07-14T07:59:56.441098Z digest=sha256:d6764c839946db308188580f625ce7e1c52edf2ce585e6874edef92d4255d42b

Observation 3e4cd782-5ba3-4906-b57b-b3866b366861 · inbound

An Exam for Active Observers cites this paper.

An Exam for Active Observers BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-01T21:12:03.439176Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T21:12:03.439176Z digest=sha256:ff78e449a8d870d704b472888c7e02bfd98bfc4938f0754f9f13e61501e6882d

Observation 968911c8-70be-4eee-8090-d09f98d0d120 · inbound

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models cites this paper.

LookME: Lookup-Based Multimodal Embeddings for Layer Injection in Vision-Language Models BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 2022

Resolution
unresolved
no resolver link, observed 2026-08-02T06:33:14.400943Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T06:33:14.400943Z digest=sha256:809273aecfdc666b0c88fad3ea254426441d21b67c2762555e10a77d916320cc

Observation fb24cba7-7b31-4c49-be5a-c6f5f1953e27 · inbound

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning cites this paper.

Trace: A Taxonomy-Guided Environment for Multidomain Visual Reasoning BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-01T11:47:23.095760Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T11:47:23.095760Z digest=sha256:3d56fdd160dae71af8f9e8c4129a23af340d31887e847fc7f0c1fb6b594da43e

Observation 933e3f93-953b-4aa7-b85e-f48e1221ae1d · inbound

Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications cites this paper.

Explainable and Resource-Efficient Spatial Reasoning in Multimodal LLMs for Decision-Critical Applications BLINK: Multimodal Large Language Models Can See but Not Perceive

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-30T12:20:51.527995Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:20:51.527995Z digest=sha256:e16779e61603f069e603912ac656ac79d1b51310a14055191d39bcf9f61c8ed9