Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-25T05:58:35.466105Z
Paper Citation Record · LEDGER
As of 10 August 2026, this Paper Citation Record lists 38 of 38 outbound references and 3 inbound Pith citation observations for arXiv:2605.22903.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-25T05:58:35.466105Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-10T06:31:04.303077+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-08-01T15:04:54.866011Z
A source-named dated measurement, never combined with another source.
Source: cited_works
38 of 38 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation bbe1f486-96d7-424c-91f0-7054527bb109 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Qwen3-vl technical report
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0abe4d4c-c03b-4302-85cd-2b32521e79bb · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Halc: Object hallucination reduc- tion via adaptive focal-contrast decoding
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 552c21fc-48ec-4058-9730-eab872861c57 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 82c008e2-0908-4703-a7c5-a2ea5468e827 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? On statistical efficiency in learning.IEEE Transactions on In- formation Theory, 67(4):2488–2506
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 507598b1-d0f3-448f-af5a-2f5b33506d7a · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Enhancing vision-language model relia- bility with uncertainty-guided dropout decoding.Advances in Neural Information Processing Systems, 38:149193– 149218
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2667e7d8-a612-4032-a7b8-f24cdd80279c · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Mme: A comprehensive evaluation benchmark for multimodal large language models
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 44229f4b-d8e8-4844-84ad-7a99734b9ae4 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Does ob- ject grounding really reduce hallucination of large vision- language models?
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 947ae1ed-30ae-483b-b353-83d142c2bb61 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Hal- lusionbench: An advanced diagnostic suite for entangled language hallucination and visual illusion in large vision- language models
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 61bd3256-36dc-4fbd-9ac0-82fbdcb23f19 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Do vision-language models really understand visual lan- guage?
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4681251b-5cfc-4e35-ad9a-11c63022dfb1 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? A survey on evaluation of multimodal large language models
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cfba36bb-ffba-4f15-8cf3-b8625cdd2523 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Robustifying vision-language models via dynamic token reweighting
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation cffb5a42-1fe7-4f34-a153-2fab6f2df224 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? A comprehensive analysis for visual object hallucination in large vision-language mod- els
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 4c60f163-65b7-4e8e-bc91-9aaae62c3554 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Do you see me : A mul- tidimensional benchmark for evaluating visual perception in multimodal llms
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 66a0dbf9-2dc0-4d41-9339-4c7251f2b2f7 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 2e54ad26-7c3b-4709-a535-267b303a00a7 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Halp: Detecting hallucinations in vision- language models without generating a single token
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation c5d31bb2-1eea-4429-b775-cb854953af8c · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? VLind-bench: Measuring language priors in large vision- language models
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 61a0e2ee-6037-406d-91cf-cec25196026f · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Evaluating object hallucination in large vision-language models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d9b41419-e78a-47a7-973a-0a00a3ca8605 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Text or pixels? evaluating efficiency and understanding of LLMs with vi- sual text inputs
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 9a45dd42-6685-4f8a-971f-8f440030f551 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? On the predictive power of representation dispersion in lan- guage models
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 27a0c21f-cfe2-4c96-a0ca-7d1894b3c780 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Visual instruction tuning
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 45fec4a6-30ed-4fec-bacb-23411fda4561 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Improved baselines with visual instruction tuning
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d5a0d606-df2b-4ff0-8f2f-633a89bf4bf0 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Grounding dino: Marry- ing dino with grounded pre-training for open-set object de- tection
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation ef4fa151-afe6-440d-b380-f383da9658f4 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? H-pope: Hierarchical polling-based probing evaluation of hallucinations in large vision-language models
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7e12274c-f23f-4824-8eb8-f4eb33aa2341 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Learning transferable visual models from natural language supervision
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 260af8aa-7554-4818-8223-def6d7ceeac8 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Sam 2: Segment anything in images and videos
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d93c712f-4fe3-4800-b463-ff00bc216cc2 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Object hallucination in image cap- tioning
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0daaaf25-dd19-4fb2-a546-397cc0b490e5 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? The effective rank: A mea- sure of effective dimensionality
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation bb71e7c2-0b06-4a11-9331-f8655048e47c · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? ISO-Bench: Benchmarking Multimodal Causal Reasoning in Visual-Language Models through Procedural Plans
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 02ef412c-421d-4547-8520-da619b63df75 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? A-okvqa: A benchmark for visual question answering using world knowl- edge
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 120be32e-0543-4118-9022-460690a2f274 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? From behavioral performance to internal competence: Interpreting vision-language models with vlm- lens
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b0323300-e024-4c4f-87bf-9b40bbf91ebc · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Openai gpt-5 system card
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b7f055d5-0f5d-48f1-b2d3-c4e6dc8bd7a1 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? From head to tail: Towards balanced representation in large vision-language models through adaptive data calibration
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0ad782f6-701b-4e5e-8fdb-606186261e0d · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Unresolved cited work
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation d8014817-b633-43a5-8298-515acb0150f5 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Unresolved cited work
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation b23ff460-f89e-400d-9b19-7971c18df27d · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Eyes wide shut? exploring the visual shortcomings of multimodal llms
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 146fb5d1-6a65-41a6-915a-ab71d202d2a0 · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Gomez, Lukasz Kaiser, and Illia Polosukhin
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 7c74344a-82d7-4d5a-afdc-65411650281d · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? Amber: An llm-free multi- dimensional benchmark for mllms hallucination evaluation
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation fcd32957-bc16-4b8f-9ce6-1d80b3451caa · outbound
Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision? no images
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-10T06:31:04.303077+00:00.
Observation 0cfc13e8-9803-4225-9a17-91e9ffaaf842 · inbound
Attention Without Grounding: Causal Evaluation of Visual Explanations in Medical VLMs Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation eb984bca-97a2-4abf-9791-12f2555fddcd · inbound
Visual Credit Audit for Multimodal Spatial Reasoning Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7f2c230e-7d58-4dee-82b8-89759e04746b · inbound
Visual Credit Audit for Multimodal Spatial Reasoning Seeing without Looking: Do Vision-Language Benchmarks Really Test Vision?
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.