Pith. sign in

Paper Citation Record · LEDGER

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

As of 10 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 3 inbound Pith citation observations for arXiv:2605.22903.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2605.22903 v1

Coverage vector

measured 38 of 38 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-25T05:58:35.466105Z

measured 41 of 41 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00

measured 3 of 3 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-01T15:04:54.866011Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

38 of 38 outbound references displayed

  • verified exact2
  • verified fuzzy32
  • unresolved3
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation bbe1f486-96d7-424c-91f0-7054527bb109 · outbound

This paper cites Qwen3-vl technical report.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Qwen3-vl technical report

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.414805Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:5267996d1627bd94fe9098c8587479287290eed9f8c12a8bb47411040632288a

Observation 0abe4d4c-c03b-4302-85cd-2b32521e79bb · outbound

This paper cites Halc: Object hallucination reduc- tion via adaptive focal-contrast decoding.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Halc: Object hallucination reduc- tion via adaptive focal-contrast decoding

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.417944Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:4162090c5aff34771798889e039b61f1403d18114290881b0bf61ec00903f6cd

Observation 552c21fc-48ec-4058-9730-eab872861c57 · outbound

This paper cites Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.404362Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:a4f5354f3e3a9803549b5d7c507aadc145ef7d47a749a23c565d114b5b430774

Observation 82c008e2-0908-4703-a7c5-a2ea5468e827 · outbound

This paper cites On statistical efficiency in learning.IEEE Transactions on In- formation Theory, 67(4):2488–2506.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? On statistical efficiency in learning.IEEE Transactions on In- formation Theory, 67(4):2488–2506

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.408744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:6aa61d79f4ef23d9bb823fa4266bc61160be37fa1979fc54766caa51638abe12

Observation 507598b1-d0f3-448f-af5a-2f5b33506d7a · outbound

This paper cites Enhancing vision-language model relia- bility with uncertainty-guided dropout decoding.Advances in Neural Information Processing Systems, 38:149193– 149218.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Enhancing vision-language model relia- bility with uncertainty-guided dropout decoding.Advances in Neural Information Processing Systems, 38:149193– 149218

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.430218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:7876af14f5b8f65e8ecb3896e6d87396c49578f39d9aef9ac55dc8368f2eb738

Observation 2667e7d8-a612-4032-a7b8-f24cdd80279c · outbound

This paper cites Mme: A comprehensive evaluation benchmark for multimodal large language models.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Mme: A comprehensive evaluation benchmark for multimodal large language models

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.433280Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:52e5693903fffbfda12d307fac9f07dee75c8b82b146e42723c321365ddd4ab1

Observation 44229f4b-d8e8-4844-84ad-7a99734b9ae4 · outbound

This paper cites Does ob- ject grounding really reduce hallucination of large vision- language models?.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Does ob- ject grounding really reduce hallucination of large vision- language models?

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.436022Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:e38bc6e17f9ddc4b7af6a2fc04eadca8105ca2b264f39a714b4273a34f5eb932

Observation 947ae1ed-30ae-483b-b353-83d142c2bb61 · outbound

This paper cites Hal- lusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Hal- lusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.455640Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:6a68302d2f63ede57b7b252c3d4630ed6da0df814e50a4be4f230b486f3a20b7

Observation 61bd3256-36dc-4fbd-9ac0-82fbdcb23f19 · outbound

This paper cites Do vision-language models really understand visual lan- guage?.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Do vision-language models really understand visual lan- guage?

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.366989Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:213d789526acdd3f8c4aa21808e439d66f0fb78a01a671e76ef32605e6e77a5a

Observation 4681251b-5cfc-4e35-ad9a-11c63022dfb1 · outbound

This paper cites A survey on evaluation of multimodal large language models.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? A survey on evaluation of multimodal large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.360769Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:e01342be7bf616fdbeec2cc04a75822f10ed011545616bca12632bb30b2f0f70

Observation cfba36bb-ffba-4f15-8cf3-b8625cdd2523 · outbound

This paper cites Robustifying vision-language models via dynamic token reweighting.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Robustifying vision-language models via dynamic token reweighting

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:00:23.370424Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:fed7528cf6ddfe3c6bf0ac2844989d915fb49ad362d3fbebf7e5daf36f2e2ad7

Observation cffb5a42-1fe7-4f34-a153-2fab6f2df224 · outbound

This paper cites A comprehensive analysis for visual object hallucination in large vision-language mod- els.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? A comprehensive analysis for visual object hallucination in large vision-language mod- els

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.351708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:428b70c98e22861f25c024ca78964f67f7729bfecdd0d1b3bed0e780001f5d39

Observation 4c60f163-65b7-4e8e-bc91-9aaae62c3554 · outbound

This paper cites Do you see me : A mul- tidimensional benchmark for evaluating visual perception in multimodal llms.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Do you see me : A mul- tidimensional benchmark for evaluating visual perception in multimodal llms

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.354421Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:1cb6fa739f1c4746b9ec157c02ad38966a28c78d4d92835c33b4e4ee7b41c7af

Observation 66a0dbf9-2dc0-4d41-9339-4c7251f2b2f7 · outbound

This paper cites an unresolved cited work.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-05-25T06:00:24.370018Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:69a7ffa5aa958a85f43fa2eb2e327136aebff3f4727d29e2bb81a80e35352662

Observation 2e54ad26-7c3b-4709-a535-267b303a00a7 · outbound

This paper cites Halp: Detecting hallucinations in vision- language models without generating a single token.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Halp: Detecting hallucinations in vision- language models without generating a single token

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.372796Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:5168b3b5d60f329a17c4f97c5b4cd2280dd528bc5a79b974f859e2c839e8c56f

Observation c5d31bb2-1eea-4429-b775-cb854953af8c · outbound

This paper cites VLind-bench: Measuring language priors in large vision- language models.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? VLind-bench: Measuring language priors in large vision- language models

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.396618Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:39688a61041e5b9b6c1b6df87c4c28bc599556e4bfb1613573c9097011bd40c5

Observation 61a0e2ee-6037-406d-91cf-cec25196026f · outbound

This paper cites Evaluating object hallucination in large vision-language models.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Evaluating object hallucination in large vision-language models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.441466Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:18822892e965bc5bf57c07455466cde637a180e9530def2dccdcc3774f62da93

Observation d9b41419-e78a-47a7-973a-0a00a3ca8605 · outbound

This paper cites Text or pixels? evaluating efficiency and understanding of LLMs with vi- sual text inputs.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Text or pixels? evaluating efficiency and understanding of LLMs with vi- sual text inputs

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.450173Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:a6760634c05661ffb802925f691c7ebfa5653b3033fa625db1b12bd757b32bd0

Observation 9a45dd42-6685-4f8a-971f-8f440030f551 · outbound

This paper cites On the predictive power of representation dispersion in lan- guage models.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? On the predictive power of representation dispersion in lan- guage models

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.438835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:889ea92f2e5ed5f8ce61ae5bea54df2ec0bcad639ffe9e624e390e7fee066252

Observation 27a0c21f-cfe2-4c96-a0ca-7d1894b3c780 · outbound

This paper cites Visual instruction tuning.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Visual instruction tuning

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.452744Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:f3f321796e6e0a3fa51f0cb206a2749ba5363c633352c5b1e55fc75bb80fbdca

Observation 45fec4a6-30ed-4fec-bacb-23411fda4561 · outbound

This paper cites Improved baselines with visual instruction tuning.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Improved baselines with visual instruction tuning

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.399786Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:7509edd3b458c6ca54e0ede196602e0c22d3d7806e68bd6f5f8045e53a24e8bc

Observation d5a0d606-df2b-4ff0-8f2f-633a89bf4bf0 · outbound

This paper cites Grounding dino: Marry- ing dino with grounded pre-training for open-set object de- tection.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Grounding dino: Marry- ing dino with grounded pre-training for open-set object de- tection

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.444251Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:9da0ad6ea909c7882b03cb809dd6174bbf5bb36725e322e06b8c8b2262aad35d

Observation ef4fa151-afe6-440d-b380-f383da9658f4 · outbound

This paper cites H-pope: Hierarchical polling-based probing evaluation of hallucinations in large vision-language models.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? H-pope: Hierarchical polling-based probing evaluation of hallucinations in large vision-language models

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.378422Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:59a5bd00e158b77a262ec4134b7c3c43e8f0ad4a3d0f4e55a993fd7c462771e0

Observation 7e12274c-f23f-4824-8eb8-f4eb33aa2341 · outbound

This paper cites Learning transferable visual models from natural language supervision.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Learning transferable visual models from natural language supervision

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.381260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:eb1ec1e3f8198ff3bfb7f38fec72c702e9ff6db0c1a96c4b7371f43bca415c76

Observation 260af8aa-7554-4818-8223-def6d7ceeac8 · outbound

This paper cites Sam 2: Segment anything in images and videos.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Sam 2: Segment anything in images and videos

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.384382Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:2bf45c05872f2e403865a8c12ef5e23242416a03198dd5185afb2f881bf5668c

Observation d93c712f-4fe3-4800-b463-ff00bc216cc2 · outbound

This paper cites Object hallucination in image cap- tioning.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Object hallucination in image cap- tioning

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.387523Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:443fa8ab366b4c4a28b8bf61e9b3ea415ee80e4381b8c687d655b12c78fe5623

Observation 0daaaf25-dd19-4fb2-a546-397cc0b490e5 · outbound

This paper cites The effective rank: A mea- sure of effective dimensionality.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? The effective rank: A mea- sure of effective dimensionality

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.390543Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:2abd655fbcd39f399a748f32a2644d6587079bf2acc3170c1aa4f7841a680ad8

Observation bb71e7c2-0b06-4a11-9331-f8655048e47c · outbound

This paper cites ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans

Reference 28

Resolution
verified exact
arxiv_id, observed 2026-05-25T06:00:23.374343Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:b4108e66cbe6b21d3a9fa3c9d97662ef8d3a76cef287953d858284c8a09a54ed

Observation 02ef412c-421d-4547-8520-da619b63df75 · outbound

This paper cites A-okvqa: A benchmark for visual question answering using world knowl- edge.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? A-okvqa: A benchmark for visual question answering using world knowl- edge

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.363858Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:229ec24521e8cfde242b961fda96ca0fa7f9706d99ab6d1c1d9c5f673daac7d4

Observation 120be32e-0543-4118-9022-460690a2f274 · outbound

This paper cites From behavioral performance to internal competence: Interpreting vision-language models with vlm- lens.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? From behavioral performance to internal competence: Interpreting vision-language models with vlm- lens

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.357415Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:990cd85a1f9d3a4ca6b74e934383255f0c8125d1210dee62062ce85a0851e739

Observation b0323300-e024-4c4f-87bf-9b40bbf91ebc · outbound

This paper cites Openai gpt-5 system card.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Openai gpt-5 system card

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.375826Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:f785d0c26dd33b93688d3df4dfb099a0c419d6270d3e6a6adfa3773b400ea040

Observation b7f055d5-0f5d-48f1-b2d3-c4e6dc8bd7a1 · outbound

This paper cites From head to tail: Towards balanced representation in large vision-language models through adaptive data calibration.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? From head to tail: Towards balanced representation in large vision-language models through adaptive data calibration

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.341090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:9001253229e3df46d274b4a4f2b8e06c68ec121d47d8af4d2ac44560aaf34183

Observation 0ad782f6-701b-4e5e-8fdb-606186261e0d · outbound

This paper cites an unresolved cited work.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Unresolved cited work

Reference 33

Resolution
unresolved
raw_fallback, observed 2026-05-25T06:00:24.343505Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:aa5e112688197b3e4cbb8a1285ec8260a2cd2d0ed3efae763810365df808d4b1

Observation d8014817-b633-43a5-8298-515acb0150f5 · outbound

This paper cites an unresolved cited work.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Unresolved cited work

Reference 34

Resolution
unresolved
raw_fallback, observed 2026-05-25T06:00:24.346183Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:29671d8ccf771a3900eecdf14d3d5a51ba19ac2b64ab0a48c0b77f2922c1b041

Observation b23ff460-f89e-400d-9b19-7971c18df27d · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.348881Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:61194c94236a66c8b176a81b70471eaf4fe244c5b86cb209f18eab9f6ba4b930

Observation 146fb5d1-6a65-41a6-915a-ab71d202d2a0 · outbound

This paper cites Gomez, Lukasz Kaiser, and Illia Polosukhin.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Gomez, Lukasz Kaiser, and Illia Polosukhin

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.393477Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:0cc0a76fa5698060166987cefb3b1126ab0a95a2d37876b19df5392fcfe04af6

Observation 7c74344a-82d7-4d5a-afdc-65411650281d · outbound

This paper cites Amber: An llm-free multi- dimensional benchmark for mllms hallucination evaluation.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Amber: An llm-free multi- dimensional benchmark for mllms hallucination evaluation

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-05-25T06:00:24.458707Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:39cca6636e56bf19cd65e2e5fe7b7089adc028da86992431d0b7050d55414248

Observation fcd32957-bc16-4b8f-9ce6-1d80b3451caa · outbound

This paper cites no images.

Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? no images

Reference 38

Resolution
malformed identifier
raw_fallback, observed 2026-05-25T06:00:24.447119Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.

source=pdf_text observed=2026-05-25T05:58:35.466105Z digest=sha256:f69341f780f07d67ba1123631dc40d8cdef8eac1d5a19e4f66e004751fc1fe00

Pith citing papers

Observation 0cfc13e8-9803-4225-9a17-91e9ffaaf842 · inbound

Attention Without Grounding: Causal Evaluation of Visual Explanations in Medical VLMs cites this paper.

Attention Without Grounding: Causal Evaluation of Visual Explanations in Medical VLMs Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-01T15:04:54.866011Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T15:04:54.866011Z digest=sha256:7ebc78f18fc0e6fa1da7fc30aa441850969943f812cff33f1bf28e862de35d1d

Observation eb984bca-97a2-4abf-9791-12f2555fddcd · inbound

Visual Credit Audit for Multimodal Spatial Reasoning cites this paper.

Visual Credit Audit for Multimodal Spatial Reasoning Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-30T12:08:02.370883Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-30T12:08:02.370883Z digest=sha256:87142295c3dd906f88594711ff723f6deb239d3964caa9adc3bb83dbb150daa4

Observation 7f2c230e-7d58-4dee-82b8-89759e04746b · inbound

Visual Credit Audit for Multimodal Spatial Reasoning cites this paper.

Visual Credit Audit for Multimodal Spatial Reasoning Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-01T10:15:14.125422Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-01T10:15:14.125422Z digest=sha256:f6903c5a223a6f31f649a766af63bac87aa42bcaaf1d9ec233d71afb1c37fbac