Pith. sign in

Paper Citation Record · LEDGER

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions

As of 4 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 2 inbound Pith citation observations for arXiv:2511.17722.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.17722 v3

Coverage vector

measured 29 of 29 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-05-17T20:05:20.393575Z

measured 31 of 31 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-20T12:19:42.722173Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: pith, observed 2026-05-20T12:23:17.003497Z

Reference resolution

29 of 29 outbound references displayed

  • verified exact6
  • verified fuzzy18
  • unresolved4
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch1

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d50246e3-fae2-4136-a91d-bdb213db25b0 · outbound

This paper cites [de— re] constructing vlms’ reasoning in counting.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions [de— re] constructing vlms’ reasoning in counting

Reference 1

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:10:11.255787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:ffb967752e68540ec39dcdefee303840bf2f845902192e94c5099af2be22d944

Observation a26fea90-62e4-48cb-9b64-df5d58b29896 · outbound

This paper cites Countgd: Multi-modal open-world counting.Advances in Neural Information Processing Systems, 37:48810–48837.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Countgd: Multi-modal open-world counting.Advances in Neural Information Processing Systems, 37:48810–48837

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:12:05.027611Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:87742283d273dad2e60bb7c8ee24d6f59c94e2242164a18020f2350503fbb25f

Observation d89c2326-4e53-41c0-9c4e-8c0f05130b4b · outbound

This paper cites Mitigating object hallucinations in large vision- language models with assembly of global and local attention.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Mitigating object hallucinations in large vision- language models with assembly of global and local attention

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:12:04.990981Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:38d006a7ea7c4cac4c9dc31cd78e06f732973992cfde9d5acfa7d48146811fa9

Observation f73601dd-ed6a-42dd-a0b8-1a1d3ba39c4b · outbound

This paper cites Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:12:04.985016Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:62fd8e4708254cfde6018dcbd8895d806afd86b7b66610af521b66bb9b9a7b20

Observation 9774e28b-d991-44de-a42e-74e0be0cce95 · outbound

This paper cites Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:12:05.022671Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:7c0a281a395bd3bf81e287ef7c9f0e7d7c6518da15482da3e340ae25dd5bcd91

Observation 6212dcc5-5390-4bb2-8213-47a59f953785 · outbound

This paper cites Probing the Visualization Literacy of Vision Language Models: the Good, the Bad, and the Ugly.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Probing the Visualization Literacy of Vision Language Models: the Good, the Bad, and the Ugly

Reference 6

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:10:11.270170Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:7a493886da44955bfe6cb3241923fa6dd8dd03eb8b590b529ac8cdf105df2dee

Observation a9ade269-d54a-4e96-8ffc-b43909c3e09f · outbound

This paper cites Google deepmind: Gemini 2.5 pro, 2025.https: //deepmind.google/models/gemini/pro/.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Google deepmind: Gemini 2.5 pro, 2025.https: //deepmind.google/models/gemini/pro/

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:12:05.018425Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:e3c926e27273e13aa44940152cecfdbba35b4e4f6a5f96794f5ac04026673e82

Observation 47c2a054-1f45-4bb9-8f8b-b7dea0792179 · outbound

This paper cites Your vision-language model can’t even count to 20: Exposing the failures of vlms in compositional counting.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Your vision-language model can’t even count to 20: Exposing the failures of vlms in compositional counting

Reference 8

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:10:11.274767Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:9c7d5c2c6a5acc253ca6216a9758e0b3820399ead142661e3445cfc73cec3e85

Observation 52fd7b0c-2cf3-4531-8914-32d5e35dfeac · outbound

This paper cites Mask r-cnn.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Mask r-cnn

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:12:05.024834Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:17e437057e6d36ee60c59091075c8d22953927df47e4b70a22b4b3c26ff6be70

Observation bb536c98-020c-4ac6-8c8b-7f2f360deeef · outbound

This paper cites Do Vision-Language Models Really Understand Visual Language?.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Do Vision-Language Models Really Understand Visual Language?

Reference 10

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:10:11.284722Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:e3e79ad14f74a9f62e30b7cf12f4c62a4d77aa3ea7b0afd3cc8eaf787c1a6af0

Observation ad1050a2-488b-44a8-b789-af7e8bdb90ba · outbound

This paper cites Point segment and count: A gener- alized framework for object counting.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Point segment and count: A gener- alized framework for object counting

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:12:05.030162Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:ff304b03c1760c67883991fa972fc0e4ba125742eaa739435cdfdd75720bb974

Observation 36680381-e702-4798-bef3-21fc87835478 · outbound

This paper cites See What You Are Told: Visual Attention Sink in Large Multimodal Models.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions See What You Are Told: Visual Attention Sink in Large Multimodal Models

Reference 12

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:10:11.265815Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:a7074cdfa80e721b1884304c8402f219aff5104521fa5926f5b2d631d5011d01

Observation 43db3c0b-0df8-4f4d-8b58-bd6c4fc60dd3 · outbound

This paper cites Segment any- thing.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Segment any- thing

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:12:05.032339Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:bc6510345ad265613f0bb7030033f1c933a1cfc1f2126b59c36abaeff840ecdd

Observation 83f6e0c4-13b4-418b-8c0a-1db9950c9f0a · outbound

This paper cites VLind-Bench: Measuring Language Priors in Large Vision-Language Models.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions VLind-Bench: Measuring Language Priors in Large Vision-Language Models

Reference 14

Resolution
metadata mismatch
arxiv_id, observed 2026-05-17T20:10:11.260463Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:dc226e51ebaeca32fada4af7e00b742e4b0a3db9fee0ae21b6817094ba579e82

Observation bbff3afe-903e-45b8-8e47-1062697a8739 · outbound

This paper cites Vhelm: A holistic evaluation of vision language models.Advances in Neural Information Processing Systems, 37:140632–140666.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Vhelm: A holistic evaluation of vision language models.Advances in Neural Information Processing Systems, 37:140632–140666

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:12:05.009905Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:c9e49ccbb908155dccffd7532241604c439cb52b60ae794d5881f9cad471bf9c

Observation 624ba155-1d3c-4e23-86fc-9935c6a4592a · outbound

This paper cites Open ai: Introducing openai o3 and o4-mini.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Open ai: Introducing openai o3 and o4-mini

Reference 16

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:12:05.011895Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:19b8a0acc8023d7e028732e5966d9b7a83e2a01b69c767f2242e29483b49fc49

Observation 2735e050-9190-458d-81eb-784192dcac4b · outbound

This paper cites Crowd- diff: Multi-hypothesis crowd density estimation using dif- fusion models.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Crowd- diff: Multi-hypothesis crowd density estimation using dif- fusion models

Reference 17

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:12:05.014360Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:3b049c8b0bf027a9b3fa91f9397a1b460ba88248718c63d17fa44d59d68fa2b7

Observation ad766355-5474-4800-a485-569676b55038 · outbound

This paper cites Vision-language foundation models for medical imag- ing: a review of current practices and innovations.Biomedi- cal Engineering Letters, pages 1–22.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Vision-language foundation models for medical imag- ing: a review of current practices and innovations.Biomedi- cal Engineering Letters, pages 1–22

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:12:05.014782Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:15f24bed477c355d2a2b8de78f7a0ce90e058ddfd06e115c797662a1dfc76dc4

Observation 21312160-65b6-4ecf-b4af-057621342db8 · outbound

This paper cites Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey

Reference 19

Resolution
verified exact
arxiv_id, observed 2026-05-17T20:28:16.822386Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:021ccca7881ab4197101a134699e80b5299d4c91c657f45b5be83371d13bd791

Observation 581acd19-16b8-47c0-9fa5-044bef8a5b6f · outbound

This paper cites an unresolved cited work.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Unresolved cited work

Reference 20

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:12:05.016160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:7ea46d11e5718f2790615dd9a6e12f4b03539c7f0e3aa00ae0ece357c95d68c1

Observation ae249483-3f8b-4838-b9e5-4756a27e24e8 · outbound

This paper cites Qwen2.5-vl.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Qwen2.5-vl

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:12:05.020643Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:7bac1660f4617997db0bb6eaaabb906c258f3ce3859a78862555703001b17318

Observation 1b0deaa2-9c69-41e9-a477-ddc9084260cf · outbound

This paper cites Vi- sion language models are biased.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Vi- sion language models are biased

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:12:05.003503Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:b2f23f75d2b79bcdaf1da97ddea4adc6dfa06e7195a1fcabb7bb1f2263fa538b

Observation 3e68b964-198f-4516-835f-e2207cadee5a · outbound

This paper cites an unresolved cited work.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Unresolved cited work

Reference 23

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:12:04.995735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:bb1185b776d880411b4e3d8fe3a467b9a322992cf78f52cb234e731399e58f2b

Observation 29b077c0-1770-4c3a-9da3-ccd5da4963ab · outbound

This paper cites an unresolved cited work.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Unresolved cited work

Reference 24

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:12:05.004377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:9307e699552bcc754f60dfbe7967c6a27afddda67f72ace9d995d74a3934a7bc

Observation 5a9901a5-6a9c-4cd3-a0f8-a11703245514 · outbound

This paper cites Count the number of objects in this image. Answer the count within curly brackets, eg.{10}.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Count the number of objects in this image. Answer the count within curly brackets, eg.{10}

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:12:05.002187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:effe2563244662f7fc8b6d33bee767660a86960883710a1d8e7c84d15b2168fd

Observation 75b6cfe4-14b3-4cc8-9b37-0f85002bd795 · outbound

This paper cites Example images for theObjectcategory,Colorpattern, showing different object colors.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Example images for theObjectcategory,Colorpattern, showing different object colors

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:12:05.034554Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:1bc0d31d8d4ef65d15f0e7a524077d9686fc709091c8968cc0424c54be4e3fc5

Observation e33c48ad-429c-4133-b0ed-cd128b19bb77 · outbound

This paper cites circles”(as default in color experiment), “squares.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions circles”(as default in color experiment), “squares

Reference 27

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:12:04.991241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:5424e1337174bff19ab99495cb5ae64361b0d741845d5f127753d6819294c3a3

Observation 6740fb6f-adcd-487c-8d1c-78cba9722958 · outbound

This paper cites an unresolved cited work.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Unresolved cited work

Reference 28

Resolution
unresolved
raw_fallback, observed 2026-05-17T20:12:05.007379Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:d298c85e576fe3b2b18592365c42ef5613d3421d4e6b80a5d84de3da060e3a0e

Observation 24a5578e-36e6-437c-8cfd-45242978a334 · outbound

This paper cites blue-green.

Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions blue-green

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-05-17T20:12:05.005528Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-17T20:05:20.393575Z digest=sha256:f9c1980314f5387c25f4ec7f5d0ac9a84911a9094eef5b581f7ebc572a360f87

Pith citing papers

Observation e88ce197-84e4-4fd7-9f42-1bee320535ed · inbound

PushupBench: Your VLM is not good at counting pushups cites this paper.

PushupBench: Your VLM is not good at counting pushups Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions

Reference 9

Resolution
verified exact
local_arxiv, observed 2026-05-11T20:36:08.900436Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=arxiv_source observed=2026-05-08T08:32:47.354181Z digest=sha256:15b78938529b4dff65a06764e89980275a11075e3fddee3a66358cf3c741e107

Observation bd968b84-02ec-406f-8382-f5a9d0d181df · inbound

CounterCount: A Diagnostic Framework for Counting Bias in Vision Language Models cites this paper.

CounterCount: A Diagnostic Framework for Counting Bias in Vision Language Models Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions

Reference 38

Resolution
verified exact
local_arxiv, observed 2026-05-20T12:23:17.005033Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.

source=pdf_text observed=2026-05-20T12:19:42.722173Z digest=sha256:35a208310b9988a19269d8615adb5b72df445e4d416e86b47ce4438aace1107f