Pith. sign in

Paper Citation Record · LEDGER

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs

As of 4 August 2026, this Paper Citation Record lists 30 of 30 outbound references and 0 inbound Pith citation observations for arXiv:2604.10528.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2604.10528 v5

Coverage vector

measured 30 of 30 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-07-12T22:34:11.603463Z

measured 30 of 30 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-04T06:34:03.388597+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

30 of 30 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved29
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 5b645d1a-c999-44bf-822b-432385986a75 · outbound

This paper cites Pixtral 12B.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Pixtral 12B

Reference 1

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:97fa8a94fc64077a07b041eeb74f94bc1408fa03d400cb0ade0fe8f709edce30

Observation 3e94dedf-4a22-4964-89e5-d00accf5e328 · outbound

This paper cites Ovis2.5: Structural embedding alignment for multimodal large language model, 2025.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Ovis2.5: Structural embedding alignment for multimodal large language model, 2025

Reference 2

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:2bfce0561488badf550d2bc11e56804d30a0e22e86a4dcf4052077032bc65031

Observation 13114dde-0758-4625-a9ed-9d070867ac1e · outbound

This paper cites Claude 4.5 model card, 2025.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Claude 4.5 model card, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:563c3e4b1db979414dd0f89a6ad79e3cc1d95ff2ef1a12661122d0c3904ccfe4

Observation 1f485e2b-12d1-4ba5-bb58-3637d83fb8c3 · outbound

This paper cites Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Qwen-VL: A Versatile Vision-Language Model for Understanding, Localization, Text Reading, and Beyond

Reference 4

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:274b2b7dd084565e6e57be11ad76ad97945c154a6379d54ec3d1540e9d5937b1

Observation d4c820be-b9f2-42b3-a7b5-3f95dfa0d930 · outbound

This paper cites Re:Verse - Can Your VLM Read a Manga? InProceedings of the IEEE/CVF Inter- national Conference on Computer Vision, pages 3761–3771,.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Re:Verse - Can Your VLM Read a Manga? InProceedings of the IEEE/CVF Inter- national Conference on Computer Vision, pages 3761–3771,

Reference 5

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:f823ca61305334022aadb94b247fa240e96f5583b0a3f5a618f410e9018fb8f4

Observation e4648a73-0c01-4bd0-b476-3fd020ba58a4 · outbound

This paper cites Florence-vl: Enhancing vision-language models with generative vision encoder and depth-breadth fusion, 2024.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Florence-vl: Enhancing vision-language models with generative vision encoder and depth-breadth fusion, 2024

Reference 6

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:045be5ad91c629289fcab00f556136e41245b02c7f44a633f1220b8657bbe5b0

Observation 22ea92c8-83e6-4c7b-9fe8-8949a65996a4 · outbound

This paper cites The pascal visual object classes (voc) challenge.IJCV, 2010.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs The pascal visual object classes (voc) challenge.IJCV, 2010

Reference 7

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:b72700e2233a5477f46a9aa8476d5a17a895aef8e4dcc8a8c8dd277d22c59bf3

Observation 0e01b4e4-9dc9-404c-a990-c82b9280962e · outbound

This paper cites Large-scale unsupervised semantic seg- mentation.IEEE TPAMI, 2022.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Large-scale unsupervised semantic seg- mentation.IEEE TPAMI, 2022

Reference 8

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:04aa5344b2baae055738391ac41ffbea87b87b99afebdf14a6e0750ed617059b

Observation b9c50db6-5c30-4e78-9e41-f4bc5ef33af9 · outbound

This paper cites Are vision language models texture or shape biased and can we steer them? InMMFM Workshop @ CVPR, 2024.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Are vision language models texture or shape biased and can we steer them? InMMFM Workshop @ CVPR, 2024

Reference 9

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:9fc757f9a559b9040df7595115776cbd4e34cc9ee706abfa6fdc6fb32368c3b5

Observation e750d369-54cd-49c9-b5c2-b7961b001680 · outbound

This paper cites Imagenet-trained cnns are biased to- wards texture; increasing shape bias improves accuracy and robustness.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Imagenet-trained cnns are biased to- wards texture; increasing shape bias improves accuracy and robustness

Reference 10

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:dfa401d2e43e0bda1c3f707c294aff496882566e1f0d69d99a22f2a8027ae40d

Observation cbf4aa75-f6f3-4431-9a27-f1b5f062ce50 · outbound

This paper cites Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities

Reference 11

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:907386725c5f2195c95869796e93ff00b0dfe8ed22948462aa263a8d8ba4c348

Observation 42e1bb70-3f0d-4960-847c-b999349da81d · outbound

This paper cites Paligemma: A versatile 3b vlm for transfer, 2024.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Paligemma: A versatile 3b vlm for transfer, 2024

Reference 12

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:eb701948ba4a6be88309299ef169238b8e483594dc39bf4c4cb202dd16835099

Observation 37f57ba2-bc06-40d1-bf59-235f01e1438c · outbound

This paper cites The origins and prevalence of tex- ture bias in convolutional neural networks.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs The origins and prevalence of tex- ture bias in convolutional neural networks

Reference 13

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:8541394476b65f4d7f5fd93cd4ff1ea96a743ba870c6f2a358b4eb4fd9fe8a9a

Observation 37abc0a1-da38-428a-acae-d88be9cb5122 · outbound

This paper cites Scaling up visual and vision-language rep- resentation learning with noisy text supervision.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Scaling up visual and vision-language rep- resentation learning with noisy text supervision

Reference 14

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:e685271ac73e11c631d6afb5aa9b1ab0b57a9030fc048a9b1ad88b55d4727154

Observation 8f7295d2-0886-44eb-99ec-2c8fe24ca07a · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs LLaVA-OneVision: Easy Visual Task Transfer

Reference 15

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:4321dfd07114a58847a0f801a1f97f10d13451588f3bf48dd3897e4d53b3397c

Observation 54b9f056-af28-4e13-877e-c38e5da08bd0 · outbound

This paper cites Seed-bench: Benchmarking multimodal llms with generative comprehension.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Seed-bench: Benchmarking multimodal llms with generative comprehension

Reference 16

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:701f463423ebe9a5e86f2db64b52c244863f87cd36e45217beb0423f9045aeff

Observation 0983e2ea-2123-4f9f-8a57-de8c9e4a96f2 · outbound

This paper cites Deep learning for thin object segmenta- tion.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Deep learning for thin object segmenta- tion

Reference 17

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:cafc9690d4a951c4f1f0131b09d625bf7d2884e4ae7065e26ecfaed41fa3110d

Observation 3cfa1c32-2f7e-4734-a723-77fc2ba52f66 · outbound

This paper cites Visual instruction tuning.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Visual instruction tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:3132dd7c5a5117287d8d31a1df07dce4406c0bd5ceb6bf9a199aae2f48661df2

Observation b462110a-a552-4d00-abd8-ef223407bdb6 · outbound

This paper cites Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts

Reference 19

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:84a76a66e6ccb3987ff848d43be3ab446c6dbd0b98985947a4e2df1dd40690a6

Observation 84f85996-6221-417d-9d1c-52de878c9b2a · outbound

This paper cites SmolVLM: Redefining small and efficient multimodal models.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs SmolVLM: Redefining small and efficient multimodal models

Reference 20

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:d1399f5b87c3fb857947653021a632ebf574089332773c785d47a608aec23b0c

Observation 11cd585e-75da-4f05-936a-57e66565fb4d · outbound

This paper cites Phi-3 vision 128k instruct, 2024.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Phi-3 vision 128k instruct, 2024

Reference 21

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:f45b179644670c44e1d2efdc6344051096f61697a623e6ad4c43a82eb7f2392c

Observation c078542f-b371-4b1a-87a1-b309e091b85e · outbound

This paper cites Verification Learning: Make Unsupervised Neuro-Symbolic System Feasible.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Verification Learning: Make Unsupervised Neuro-Symbolic System Feasible

Reference 22

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:5633618ff6da6ee24f127e03a12a4fbc862b612345e67a13871cb9d95ccc8ded

Observation d1d4d114-b130-4118-b605-31ac9495f4aa · outbound

This paper cites Internvl2.5 pretrained models, 2024.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Internvl2.5 pretrained models, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:b00ecdc0849a5de10cd970e352afc80289c8224533af2ece90f94c62d9d05237

Observation 80735fb8-fbac-4506-a77b-a67f5672a02a · outbound

This paper cites Robust onion: Peeling Open V ocab Object Detectors Under Noise.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Robust onion: Peeling Open V ocab Object Detectors Under Noise

Reference 24

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:2f1cb8f63b25282bad78c8a71bf3425c1e8844abffd7c6145087b476eeb3ff3a

Observation e61e22ab-d32d-4ad5-bf9f-a8be7c8c0456 · outbound

This paper cites Highly accurate dichotomous image seg- mentation.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Highly accurate dichotomous image seg- mentation

Reference 25

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:fdeaddb1790a3ba9dcba541005c531f0b49891760632739f76a34d40de859548

Observation b2fd4b0d-4233-4954-824c-007aead7330b · outbound

This paper cites Qwen2.5-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Qwen2.5-vl: Enhancing vision-language model’s perception of the world at any resolution, 2024

Reference 26

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:7947599d62f9e0b5b7a0f5521ffb9fbfa2eb81ebd245763eed405a58d50adb66

Observation e95c2ff0-0cea-424d-8ac3-90eeace477d7 · outbound

This paper cites Learning transferable visual models from natural language supervision.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Learning transferable visual models from natural language supervision

Reference 27

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:343bba8803e2eeb64fe2d8ef98e5f12d9fa2ba2eabb360d292e2461993e3b332

Observation 5babb2ba-bd7b-4109-a9a8-5cf317b53757 · outbound

This paper cites Eyes wide shut? exploring the visual shortcomings of multimodal llms.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Eyes wide shut? exploring the visual shortcomings of multimodal llms

Reference 28

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:49fab3ee17e99243334478697f0bdf03f684678d2039af656650d81b40a60cd4

Observation 62df8977-9da1-467e-8634-15088da1f8f4 · outbound

This paper cites The caltech-ucsd birds-200-2011 dataset.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs The caltech-ucsd birds-200-2011 dataset

Reference 29

Resolution
unresolved
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:3f1c5c503abb0c06f8846192341fab8e67adbb407ca19067c556b52b81d5f33f

Observation 0bce14a2-83ae-4496-8e73-f2e4c25667bb · outbound

This paper cites Who’s That Pok´emon?.

BareBones: Benchmarking Zero-Shot Geometric Comprehension in VLMs Who’s That Pok´emon?

Reference 30

Resolution
malformed identifier
no resolver link, observed 2026-07-12T22:34:11.603463Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-07-12T22:34:11.603463Z digest=sha256:309c63ab7d1c4d105692d912c6152a71c2cc343c36c2157cd5df4e9a204656dd

Pith citing papers

No inbound Pith citation observations are available.