Pith. sign in

Paper Citation Record · LEDGER

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features

As of 19 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 0 inbound Pith citation observations for arXiv:2509.08266.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2509.08266 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-04T20:58:26.760637Z

measured 25 of 25 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-19T06:32:44.657259+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy12
  • unresolved13
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 8bbe823e-462c-40e3-b69f-d81d2a79d9d7 · outbound

This paper cites Vision Language Models are Biased.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Vision Language Models are Biased

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:24.320688Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:24.320688Z digest=sha256:ec0ce46ef3af9b0d9e14765c7c68f413d55e7de09851a8ebe9670baa4b8103fd

Observation 401146e2-0882-438c-8edd-f1ed24fb398a · outbound

This paper cites Open ai: Introducing openai o3 and o4-mini, 2025.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Open ai: Introducing openai o3 and o4-mini, 2025

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:30.323971Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T20:58:24.418011Z digest=sha256:4023c852822427e4ced4c59a142dc02271d3cc566cd922b4b8cd6398ae883381

Observation 5a50d699-9317-403e-b6f0-fa7c18a6eef9 · outbound

This paper cites Google deepmind: Gemini 2.5 pro, 2025.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Google deepmind: Gemini 2.5 pro, 2025

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:30.136766Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T20:58:24.537822Z digest=sha256:0e94de0a5ce21314e2507b1d1687db6a7f50c64cb1cca414f794913288cdb1e4

Observation 9f01fb84-53a7-4d57-8174-b08f4d95fdad · outbound

This paper cites Do Vision-Language Models Really Understand Visual Language?.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Do Vision-Language Models Really Understand Visual Language?

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:24.649412Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:24.649412Z digest=sha256:ddd3ff6339c8f568d54f67536c71b317259fd80b1ae1327862db58bd925035a1

Observation 42696a3c-e460-4e02-b84b-5a5fa5eeee81 · outbound

This paper cites VLind-Bench: Measuring Language Priors in Large Vision-Language Models.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features VLind-Bench: Measuring Language Priors in Large Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:24.757401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:24.757401Z digest=sha256:2f83de18c7f7c1f289eeedacbc5668df5e1f1642f4b26a58f3ce67304a49d1fd

Observation a831130e-7f07-4a94-b749-20bb5a609c37 · outbound

This paper cites Vhelm: A holistic evaluation of vision language models.Advances in Neural Information Processing Systems, 37:140632–140666, 2024.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Vhelm: A holistic evaluation of vision language models.Advances in Neural Information Processing Systems, 37:140632–140666, 2024

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:29.845483Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T20:58:24.886900Z digest=sha256:c2c1fd3508bfaa8e1b7e1abae8a003289806493bb65e049869a64a7e81725eec

Observation 6d34adfe-0899-4e2c-9216-88ccf8b798d7 · outbound

This paper cites Probing the Visualization Literacy of Vision Language Models: the Good, the Bad, and the Ugly.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Probing the Visualization Literacy of Vision Language Models: the Good, the Bad, and the Ugly

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:24.989862Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:24.989862Z digest=sha256:e14b0d96be05758a1fc3661fdd34816cb521d162ea9e20475b641c2441dce984

Observation 15927bdb-d807-466a-911a-3ca3587833fb · outbound

This paper cites Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Don't Miss the Forest for the Trees: Attentional Vision Calibration for Large Vision Language Models

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:25.135570Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:25.135570Z digest=sha256:661f0360198e5082e2594378bc2d6a8acf6e4c5ba58e3705b437b6bd5848cd34

Observation 4b317174-849f-4293-b417-c7c08d6f86af · outbound

This paper cites Mitigating object hallucinations in large vision-language models with assembly of global and local attention.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Mitigating object hallucinations in large vision-language models with assembly of global and local attention

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:29.632342Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T20:58:25.210104Z digest=sha256:80161ae77a99d5e6513c9916a3725c524005b15e84a72afa33002884b1122479

Observation 65ddff21-ce30-45c1-9483-4fd7cbd0887b · outbound

This paper cites See What You Are Told: Visual Attention Sink in Large Multimodal Models.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:25.290792Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:25.290792Z digest=sha256:0cf096185e0c2461bcc4539f166ceb2dce9e95415eda7f4c7515e245e661eff6

Observation 088d296f-b968-48ad-bad2-d164b2ac3698 · outbound

This paper cites Qwen2.5-vl, January 2025.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Qwen2.5-vl, January 2025

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:25.372227Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:25.372227Z digest=sha256:c3536045c2a5d7eaf972aa249391ab0ec9c9a30c2a9aa0e9c4727e44a0abec26

Observation 0d57f32d-4e42-4dd6-a132-2d26bd3076d7 · outbound

This paper cites Kimi-VL Technical Report.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Kimi-VL Technical Report

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-04T20:58:25.467092Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-04T20:58:25.467092Z digest=sha256:54197ea9f93318bfa32c9b010e36c7df5bf72a4b47f34e5b53f83cd47c55799f

Observation 9dae0dc1-d113-46ce-93d8-f074cc831067 · outbound

This paper cites an unresolved cited work.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-04T20:58:29.516499Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T20:58:25.565607Z digest=sha256:647da2c66cf4b60178fc49aa1cd6a6316acce70fa173c03333652acc8078d222

Observation 82fc1942-6ff6-48f2-a64b-25d2d07aeddf · outbound

This paper cites an unresolved cited work.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Unresolved cited work

Reference 14

Resolution
unresolved
raw_fallback, observed 2026-08-04T20:58:29.272055Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T20:58:25.660579Z digest=sha256:60e83f5d98c61d481587ff1f874c41afbf90489c49ae43e67675b87f2b7eedad

Observation df664d15-dfa3-41ba-8c56-15d0995bc94c · outbound

This paper cites an unresolved cited work.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Unresolved cited work

Reference 15

Resolution
unresolved
raw_fallback, observed 2026-08-04T20:58:29.061393Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T20:58:25.723912Z digest=sha256:b0747fef38c3c0e96cd8619f6cf8457f8018318c0a84839c718f05752c3b0b90

Observation 112c2379-fe21-4f38-97ed-77e6e4658966 · outbound

This paper cites an unresolved cited work.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-04T20:58:28.837798Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T20:58:25.849248Z digest=sha256:6538c8e7cae64f4967a1d2d4f02cf6d7530b26e168cb92b239b6e7963198b33e

Observation f2bc3755-82a8-4d66-a4b7-cf976b1af856 · outbound

This paper cites We report these metrics over Qwen2.5-VL-7B, Qwen2.5-VL-32B and Kimi-VL-A3B.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features We report these metrics over Qwen2.5-VL-7B, Qwen2.5-VL-32B and Kimi-VL-A3B

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:28.617063Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T20:58:25.949159Z digest=sha256:0f14af1e3a0d97a310a26474b5536c3f7e74bb70821dc4871ea7f7c741f91cb8

Observation 1ed5b933-1c0a-4ca8-8453-642c384360be · outbound

This paper cites In Flag Stars, specifying the target object and requiring structured output substantially increases accuracy (up to 0.4; Figures 8 and 9).

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features In Flag Stars, specifying the target object and requiring structured output substantially increases accuracy (up to 0.4; Figures 8 and 9)

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:28.422583Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T20:58:26.085382Z digest=sha256:175ddf98330ac34352508738f30df53847ab8fb5a2ce7a8dd77de8bfd23e6355

Observation d64897c2-e6a6-488a-b583-2d0ead436151 · outbound

This paper cites Refer Figures 5,6 and 7.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Refer Figures 5,6 and 7

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:28.218182Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T20:58:26.172634Z digest=sha256:6f4de04ccaca912d7ea158bba1a8d17900291e5d90c893e3f54c0552d6e53ded

Observation 28fd5db4-1fd8-465e-bdae-e4042f821720 · outbound

This paper cites We can see that the proportion of attention across the same prompt for different object shapes are within a very small interval.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features We can see that the proportion of attention across the same prompt for different object shapes are within a very small interval

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:27.939333Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T20:58:26.255914Z digest=sha256:1faa2cc0fbf67cbd29caceb9760ea58fee1179b3f9bc39e562db5db673a9f914

Observation 5dd5e933-aa7c-4eab-b042-7da77c559e15 · outbound

This paper cites When the number of objects in the image is <10, the models perform relatively accurately, but counting performance becomes less accurate as we move towards the >40 bucket.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features When the number of objects in the image is <10, the models perform relatively accurately, but counting performance becomes less accurate as we move towards the >40 bucket

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:27.748866Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T20:58:26.369271Z digest=sha256:f68c55cd63be2569aa8acee61f6486fea5a1791aa482a43ebb6d39cda1bcf745

Observation eee89944-913f-45a8-8018-bd2c559f5bac · outbound

This paper cites Errors for Qwen 2.5-VL are centered mostly around negative values, meaning the model often underestimates compared to ground truth.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Errors for Qwen 2.5-VL are centered mostly around negative values, meaning the model often underestimates compared to ground truth

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:27.549741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T20:58:26.458136Z digest=sha256:552fec97468c014bd9bbb5dfa433328387a47ac3d277671362e87d616b039663

Observation 8d70d92a-852e-4597-b085-c4614e156384 · outbound

This paper cites Refer Figures 8 and 9.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Refer Figures 8 and 9

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:27.330933Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T20:58:26.564356Z digest=sha256:100a8d80bd3df74f616844fbfd2039e406b15accbb68a6122b54693fe23bdee6

Observation 675d7e65-d541-4b06-a7be-441f0b0639a3 · outbound

This paper cites an unresolved cited work.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-08-04T20:58:27.138585Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T20:58:26.675080Z digest=sha256:e09568148f57e87297daa641765d5b0384dc39b9b27e615837a08d4db07ceabf

Observation 0ffc17cc-26f5-4d94-80b4-55c3c9bb822f · outbound

This paper cites [1], when using the same prompts and data as them.

Examining Vision Language Models through Multi-dimensional Experiments with Vision and Text Features [1], when using the same prompts and data as them

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-04T20:58:27.003109Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-19T06:32:44.657259+00:00.

source=pdf_text observed=2026-08-04T20:58:26.760637Z digest=sha256:51d0714cbb6de1ea274e22a3da0289757a32393c9bb28fe8b1a05fe339607fa2

Pith citing papers

No inbound Pith citation observations are available.