Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T20:05:20.393575Z
Paper Citation Record · LEDGER
As of 4 August 2026, this Paper Citation Record lists 29 of 29 outbound references and 2 inbound Pith citation observations for arXiv:2511.17722.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-05-17T20:05:20.393575Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-03T06:30:56.289259+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-20T12:19:42.722173Z
A source-named dated measurement, never combined with another source.
Source: pith, observed 2026-05-20T12:23:17.003497Z
29 of 29 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation d50246e3-fae2-4136-a91d-bdb213db25b0 · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions [de— re] constructing vlms’ reasoning in counting
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a26fea90-62e4-48cb-9b64-df5d58b29896 · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Countgd: Multi-modal open-world counting.Advances in Neural Information Processing Systems, 37:48810–48837
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation d89c2326-4e53-41c0-9c4e-8c0f05130b4b · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Mitigating object hallucinations in large vision- language models with assembly of global and local attention
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation f73601dd-ed6a-42dd-a0b8-1a1d3ba39c4b · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Generic attention- model explainability for interpreting bi-modal and encoder- decoder transformers
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 9774e28b-d991-44de-a42e-74e0be0cce95 · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Internvl: Scaling up vision foundation mod- els and aligning for generic visual-linguistic tasks
Reference 5
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6212dcc5-5390-4bb2-8213-47a59f953785 · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Probing the Visualization Literacy of Vision Language Models: the Good, the Bad, and the Ugly
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation a9ade269-d54a-4e96-8ffc-b43909c3e09f · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Google deepmind: Gemini 2.5 pro, 2025.https: //deepmind.google/models/gemini/pro/
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 47c2a054-1f45-4bb9-8f8b-b7dea0792179 · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Your vision-language model can’t even count to 20: Exposing the failures of vlms in compositional counting
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 52fd7b0c-2cf3-4531-8914-32d5e35dfeac · outbound
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation bb536c98-020c-4ac6-8c8b-7f2f360deeef · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Do Vision-Language Models Really Understand Visual Language?
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ad1050a2-488b-44a8-b789-af7e8bdb90ba · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Point segment and count: A gener- alized framework for object counting
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 36680381-e702-4798-bef3-21fc87835478 · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions See What You Are Told: Visual Attention Sink in Large Multimodal Models
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 43db3c0b-0df8-4f4d-8b58-bd6c4fc60dd3 · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Segment any- thing
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 83f6e0c4-13b4-418b-8c0a-1db9950c9f0a · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions VLind-Bench: Measuring Language Priors in Large Vision-Language Models
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation bbff3afe-903e-45b8-8e47-1062697a8739 · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Vhelm: A holistic evaluation of vision language models.Advances in Neural Information Processing Systems, 37:140632–140666
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 624ba155-1d3c-4e23-86fc-9935c6a4592a · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Open ai: Introducing openai o3 and o4-mini
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 2735e050-9190-458d-81eb-784192dcac4b · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Crowd- diff: Multi-hypothesis crowd density estimation using dif- fusion models
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ad766355-5474-4800-a485-569676b55038 · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Vision-language foundation models for medical imag- ing: a review of current practices and innovations.Biomedi- cal Engineering Letters, pages 1–22
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 21312160-65b6-4ecf-b4af-057621342db8 · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 581acd19-16b8-47c0-9fa5-044bef8a5b6f · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Unresolved cited work
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation ae249483-3f8b-4838-b9e5-4756a27e24e8 · outbound
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 1b0deaa2-9c69-41e9-a477-ddc9084260cf · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Vi- sion language models are biased
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 3e68b964-198f-4516-835f-e2207cadee5a · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 29b077c0-1770-4c3a-9da3-ccd5da4963ab · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Unresolved cited work
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 5a9901a5-6a9c-4cd3-a0f8-a11703245514 · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Count the number of objects in this image. Answer the count within curly brackets, eg.{10}
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 75b6cfe4-14b3-4cc8-9b37-0f85002bd795 · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Example images for theObjectcategory,Colorpattern, showing different object colors
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e33c48ad-429c-4133-b0ed-cd128b19bb77 · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions circles”(as default in color experiment), “squares
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 6740fb6f-adcd-487c-8d1c-78cba9722958 · outbound
Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation 24a5578e-36e6-437c-8cfd-45242978a334 · outbound
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation e88ce197-84e4-4fd7-9f42-1bee320535ed · inbound
PushupBench: Your VLM is not good at counting pushups Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.
Observation bd968b84-02ec-406f-8382-f5a9d0d181df · inbound
CounterCount: A Diagnostic Framework for Counting Bias in Vision Language Models Can Vision-Language Models Count? A Synthetic Benchmark and Analysis of Attention-Based Interventions
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-03T06:30:56.289259+00:00.