Pith. sign in

Paper Citation Record · LEDGER

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation

As of 9 August 2026, this Paper Citation Record lists 44 of 44 outbound references and 2 inbound Pith citation observations for arXiv:2511.11440.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2511.11440 v3

Coverage vector

measured 44 of 44 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-03T22:13:55.098153Z

measured 46 of 46 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-08T06:32:00.761636+00:00

measured 2 of 2 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-05-19T17:30:08.755545Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: arxiv_reference, observed 2026-05-19T17:32:41.699954Z

Reference resolution

44 of 44 outbound references displayed

  • verified exact0
  • verified fuzzy0
  • unresolved43
  • parse uncertain0
  • malformed identifier1
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation d4ab86c6-7d9d-4d2d-ae2c-e476fbd476f0 · outbound

This paper cites Tal- lyqa: Answering complex counting questions.Proceedings of the AAAI Conference on Artificial Intelligence, 33(01): 8076–8084, 2019.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Tal- lyqa: Answering complex counting questions.Proceedings of the AAAI Conference on Artificial Intelligence, 33(01): 8076–8084, 2019

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:49.149980Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:49.149980Z digest=sha256:e22695ac73141e9776424b035d9dda87577c51b0a315031b97ec18238a63614b

Observation 02582e5a-e4c4-4345-8d4d-49ae9230a62f · outbound

This paper cites Lipton, and J.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Lipton, and J

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:49.234976Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:49.234976Z digest=sha256:f03a5c5e1b061d2777bf51fac3ed406891542d88bc3348d75a5483b4c55a450c

Observation d8004afa-709a-4d7a-95a9-ef2437d2b00d · outbound

This paper cites [De|Re] constructing VLMs’ Reasoning in Counting.arXiv preprint arXiv:2510.19555, 2025.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation [De|Re] constructing VLMs’ Reasoning in Counting.arXiv preprint arXiv:2510.19555, 2025

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:49.329477Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:49.329477Z digest=sha256:673e72601f634230030523df10c8dd7738bec318d458068df587bd2aca900b90

Observation 49d98304-ba6b-42da-b096-ce0bf554b8db · outbound

This paper cites Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Spatialvlm: Endow- ing vision-language models with spatial reasoning capabili- ties

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:49.423186Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:49.423186Z digest=sha256:3be016ce289f79c5cc238a6b28d367feeaf0853a85de4820b9ed745414794ca4

Observation 98a921ef-9889-4a38-b3c0-64dc3ac7e5dd · outbound

This paper cites Are we on the right way for eval- uating large vision-language models? InAdvances in Neural Information Processing Systems, pages 27056–27087.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Are we on the right way for eval- uating large vision-language models? InAdvances in Neural Information Processing Systems, pages 27056–27087

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:49.510602Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:49.510602Z digest=sha256:da04a048429719e14fe8ad768c8276f3899782c6864d5f831ff37a3ecb101dc1

Observation f004d32a-cb71-43e4-a4a2-45c22632dcb3 · outbound

This paper cites Why is spatial reasoning hard for VLMs? an attention mechanism perspective on focus ar- eas.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Why is spatial reasoning hard for VLMs? an attention mechanism perspective on focus ar- eas

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:49.593203Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:49.593203Z digest=sha256:3f60c89616eeaf71804ae889c81b859cbf40c4e4bb434fa3cd5a0e4be1dc613b

Observation 30066842-c3b4-4b25-bd38-658ca3442d4f · outbound

This paper cites From the least to the most: Building a plug-and-play visual rea- soner via data synthesis.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation From the least to the most: Building a plug-and-play visual rea- soner via data synthesis

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:49.659507Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:49.659507Z digest=sha256:ff8a3c70e76511552efce0af20c5bd7890bab31b8b33c274a677065955af2461

Observation 4a4601bb-0f81-46f6-9a32-b047c6c3691f · outbound

This paper cites Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:49.740903Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:49.740903Z digest=sha256:257a8bf199a5390e1fa8f67c184b4e30b7aedc2cbe68245f3764344036b58dab

Observation 3614751c-62b0-4bd3-a8d4-47687692475d · outbound

This paper cites If CLIP could talk: Understanding vision-language model representations through their preferred concept descriptions.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation If CLIP could talk: Understanding vision-language model representations through their preferred concept descriptions

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:49.836972Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:49.836972Z digest=sha256:221709a643858632d6f1c64243c3ebedf581cc9ed87efa79a9f2db9952de31ea

Observation 720e2419-d17d-4c04-81da-ba5ae9c6d8a4 · outbound

This paper cites Hidden in plain sight: VLMs overlook their visual rep- resentations.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Hidden in plain sight: VLMs overlook their visual rep- resentations

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:49.927559Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:49.927559Z digest=sha256:a1a719aa6d6cbcedcd35f8bf0ff6df183bd596acb4ddd9d93ab6ff70e20575f2

Observation 1d3aaa6a-42b4-40a8-9de3-9b207f987cde · outbound

This paper cites Generate then select: Open- ended visual question answering guided by world knowl- edge.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Generate then select: Open- ended visual question answering guided by world knowl- edge

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:50.015009Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:50.015009Z digest=sha256:012aef33158c31cfdfc8b6777f9e6622b6a3876f8fbe5dc1563e8f197d8313e3

Observation 316e89e8-cee1-4927-b25c-92c6f6e17315 · outbound

This paper cites Smith, Wei-Chiu Ma, and Ranjay Krishna.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Smith, Wei-Chiu Ma, and Ranjay Krishna

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:50.152594Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:50.152594Z digest=sha256:fa4772d540a399d288254c0e943ac6e654c479064560b87ab936f5381e54aafd

Observation 36a480b3-ba6c-4c06-a795-96bf70deea0c · outbound

This paper cites G-LLaV A: Solving geometric problem with multi-modal large language model.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation G-LLaV A: Solving geometric problem with multi-modal large language model

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:50.252377Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:50.252377Z digest=sha256:8913f6df9a83021c124f0aa3a2361d40a18f2cd7cce16418f83fb72350000c11

Observation b28128d8-eb77-4543-8002-e1fe438d4e46 · outbound

This paper cites Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Making the v in vqa matter: Elevating the role of image understanding in visual question answer- ing

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:50.436120Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:50.436120Z digest=sha256:d37a5b20efd8c7e490cd0069bbb1f905281d1f034d4e7617ff14c43c3325cada

Observation 50907dbb-5bb4-4fe5-bb17-07054a6f47a1 · outbound

This paper cites LoRA: Low-rank adaptation of large language models.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation LoRA: Low-rank adaptation of large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:50.555884Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:50.555884Z digest=sha256:9e191196d7673202803de95720e0349fb70b28085189cd7b5704e0b91af1fd7e

Observation 6b00ee62-87f2-425e-bd20-bfb794e9f350 · outbound

This paper cites Lawrence Zitnick, and Ross Girshick.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Lawrence Zitnick, and Ross Girshick

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:50.749285Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:50.749285Z digest=sha256:b1921b91f8487dfbcb7c7148dd16be87481a0a764e3bc60f08af97d5ac96f28c

Observation d022c669-b391-4b87-8b80-135ee6f0122f · outbound

This paper cites What’s “up” with vision-language models? investigating their strug- gle with spatial reasoning.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation What’s “up” with vision-language models? investigating their strug- gle with spatial reasoning

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:50.897442Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:50.897442Z digest=sha256:f60cf37bb2262a9b7e1ae9908cd761eacb22c8eab849f1b49aaea74b91ce4bf6

Observation e628b165-32cd-4c05-ad72-fb8f4a817859 · outbound

This paper cites an unresolved cited work.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Unresolved cited work

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:51.059982Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:51.059982Z digest=sha256:c73f244061410f64fdd58dcbde97d27ab818ed21bc7ac4fd71c298fcc370e200

Observation 8020ff24-bcb4-46d0-8df2-9d2043873a1f · outbound

This paper cites Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Berg, Wan-Yen Lo, Piotr Dollar, and Ross Girshick

Reference 19

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:51.225520Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:51.225520Z digest=sha256:5ea1ba2e214a736d7f617e54c85607649f2a921ea2305c78f8b5210a2c5ed1fa

Observation a0e4f4e4-b149-4afb-9da9-a5d6349e233e · outbound

This paper cites Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123(1):32–73, 2017.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Visual genome: Connecting language and vision using crowdsourced dense image annotations.International journal of computer vision, 123(1):32–73, 2017

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:51.406864Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:51.406864Z digest=sha256:ecfbcf49ba9d08f6fde5c89ab9a834642c985aa21a627e40a70d4e070386d739

Observation b953bb44-b564-4225-8daa-40dc5e8694a2 · outbound

This paper cites Enhancing vision-language com- positional understanding with multimodal synthetic data.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Enhancing vision-language com- positional understanding with multimodal synthetic data

Reference 21

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:51.611263Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:51.611263Z digest=sha256:f74ed9f91f93a2d1188564efcdf91a9514a64122c96bfa26c4508c2e5c792683

Observation 17be5194-a0bf-4c10-97c9-e21c76d84c96 · outbound

This paper cites Lawrence Zitnick.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Lawrence Zitnick

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:51.817401Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:51.817401Z digest=sha256:405f03104b80fda794c459df5b494707b3b81103b5a998e3dbc46a591300d290

Observation e780ebc2-2924-40b3-87b8-7aec16baef4a · outbound

This paper cites Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Llava-next: Im- proved reasoning, ocr, and world knowledge, 2024

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:52.176932Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:52.176932Z digest=sha256:683a0243b46cd397bb04bc9924398d06fc0baaa0ad5e082905c082c8900ae195

Observation ba066641-729c-4b3d-b6c6-854f6ebf7940 · outbound

This paper cites Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Enhancing Spatial Reasoning in Multimodal Large Language Models through Reasoning-based Segmentation

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:52.383043Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:52.383043Z digest=sha256:ab5e80b154415d045e37ef9f6329f55ea55c91cf0d75ad2c4ec48ceed787881b

Observation bef716b4-6525-473c-848e-9e9c7f6afce5 · outbound

This paper cites SpaRE: Enhancing spatial reasoning in vision-language models with synthetic data.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation SpaRE: Enhancing spatial reasoning in vision-language models with synthetic data

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:52.530421Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:52.530421Z digest=sha256:94cfdeec8ec8bfedfc48d499d8df85be3162e01ade82853bd323d9b42bf7c3fa

Observation 62a8903a-3fed-4c1e-a9a8-2600dd6c57ac · outbound

This paper cites Teaching clip to count to ten.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Teaching clip to count to ten

Reference 26

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:52.658358Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:52.658358Z digest=sha256:005f6021d799d401cab45c2fac633ac6298b1082fe3e03fe9775c30ce153db4a

Observation 609479cc-f72d-45d8-91d8-50c20945268e · outbound

This paper cites Synthetic visual genome.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Synthetic visual genome

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:52.835484Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:52.835484Z digest=sha256:6a0f4efa6e0242da8eb3d2bb057a7c81cad15578d7fa67ceb7b9dfb7703f63b4

Observation ddea47cf-f4be-4333-b97e-6c8581d64af2 · outbound

This paper cites Synthesize diagnose and optimize: Towards fine- grained vision-language understanding.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Synthesize diagnose and optimize: Towards fine- grained vision-language understanding

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:52.981256Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:52.981256Z digest=sha256:dd4ede31a84b6f9cd75dde7af4607d4834248cd5e555c330ea88b3495573914a

Observation f79ce9f0-e1da-4f9b-8dd0-f6c83610665d · outbound

This paper cites Learning transferable visual models from natural language supervision.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Learning transferable visual models from natural language supervision

Reference 29

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:53.112551Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:53.112551Z digest=sha256:906ecc2d26ed0debfcb26dab83e1971d1c1e2c9370bfc1f531cc35145a319ab1

Observation dd9b9521-04ab-4880-bf3a-be1bf17acade · outbound

This paper cites Vision language models are blind.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Vision language models are blind

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:53.306094Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:53.306094Z digest=sha256:2c31ef1154da74396b389133202441e43ed55fc1fcca2ea61a57ad6a3a8d69ea

Observation 0b0c52f2-e188-4d1e-a71f-cdba2dedfb01 · outbound

This paper cites CIVET: Systematic evaluation of understanding in VLMs.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation CIVET: Systematic evaluation of understanding in VLMs

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:53.436110Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:53.436110Z digest=sha256:7dd891ebf5b3f8b73a52975165dbfe838f01b3423f63d7509330314762bd6fa7

Observation 6bfbbfa6-8cf2-4054-aaf3-1c3f674bef12 · outbound

This paper cites Sophia Koepke, Oriol Vinyals, Cordelia Schmid, and Zeynep Akata.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Sophia Koepke, Oriol Vinyals, Cordelia Schmid, and Zeynep Akata

Reference 32

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:53.674610Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:53.674610Z digest=sha256:db6a2a038772dd3dec4b38a233ad9d42edfbce1d48f91e82b0158ef027bf5986

Observation 1ef3f6a8-c342-4a3c-a0c4-6caabc9abc86 · outbound

This paper cites Forgotten polygons: Multimodal large language models are shape-blind.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Forgotten polygons: Multimodal large language models are shape-blind

Reference 33

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:53.825355Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:53.825355Z digest=sha256:02ada15a5573bcd9c162596034e05188851e944d923ef817fcda88890a8be003

Observation 11b07d57-5378-45c0-a173-ce0891d4e4f8 · outbound

This paper cites Laion- 400m: Open dataset of clip-filtered 400 million image-text pairs.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Laion- 400m: Open dataset of clip-filtered 400 million image-text pairs

Reference 34

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:53.981048Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:53.981048Z digest=sha256:9e814ee66b00afe495e1ebedf87cd28792080834975302e2469b983f5ce57f6a

Observation 2320c487-df7c-4cba-8142-bdcacd3e7bf0 · outbound

This paper cites Math- LLaV A: Bootstrapping mathematical reasoning for multi- modal large language models.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Math- LLaV A: Bootstrapping mathematical reasoning for multi- modal large language models

Reference 35

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:54.116549Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:54.116549Z digest=sha256:27d8290029d65f137a7ddd21d215ad8d0f1d7057dcaa4aa5feb71f343b8d456f

Observation fa8f5cd4-3ebe-4dc0-9ad2-8ed6915c1028 · outbound

This paper cites Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution

Reference 36

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:54.235480Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:54.235480Z digest=sha256:9e0b092e28480800cd1b7d34773f2a8fa6522b57d10035d1e13aed4ff1363056

Observation dd8d48ed-3b35-4882-8a1a-5e3afcabb0ed · outbound

This paper cites Embodied scene understanding for vi- sion language models via metavqa.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Embodied scene understanding for vi- sion language models via metavqa

Reference 37

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:54.405053Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:54.405053Z digest=sha256:1312fdb37ced67f5f911da46533abd5a3e32d3d4e51248fb48ced96e38860046

Observation adfc1cb8-958f-414f-b870-2f35faf6af75 · outbound

This paper cites Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Mmmu: A massive multi-discipline multimodal understand- ing and reasoning benchmark for expert agi

Reference 38

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:54.522472Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:54.522472Z digest=sha256:e0d56f91e3cf792c0f8fdb73e67b1ecafe2196e6ef0fbbda400ed93b25b5bb19

Observation e9166c7c-6580-4753-9070-f4263085a136 · outbound

This paper cites When and why vision- language models behave like bags-of-words, and what to 10 do about it? InThe Eleventh International Conference on Learning Representations, 2023.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation When and why vision- language models behave like bags-of-words, and what to 10 do about it? InThe Eleventh International Conference on Learning Representations, 2023

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:54.693069Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:54.693069Z digest=sha256:e0acfd9e2ac9cbe4f3801571749969f1a1a7616f2e267b9b841e85c21af6b21a

Observation 7713cc03-3acf-4d1e-8f2b-5bcdba030cfd · outbound

This paper cites Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InComputer Vision – ECCV 2024, pages 169–186, Cham, 2025.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Mathverse: Does your multi-modal llm truly see the diagrams in visual math problems? InComputer Vision – ECCV 2024, pages 169–186, Cham, 2025

Reference 40

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:54.847296Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:54.847296Z digest=sha256:fa511a31e26bfa35c9890855f94b94004f168bedb92f20d0b0df7cd92a129854

Observation a109cb86-f302-4c3b-ac49-bc49a1cf02c6 · outbound

This paper cites MA VIS: Mathe- matical visual instruction tuning with an automatic data en- gine.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation MA VIS: Mathe- matical visual instruction tuning with an automatic data en- gine

Reference 41

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:54.971591Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:54.971591Z digest=sha256:9aa3e6402a3ae63e3d17288f0aeba0b0c2f899f054bd51f562dc527ac2502ffc

Observation fe6a5bae-47e9-4d1d-bcec-172479e0bff2 · outbound

This paper cites Answer with as few words as possible.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Answer with as few words as possible

Reference 42

Resolution
malformed identifier
no resolver link, observed 2026-08-03T22:13:55.098153Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:55.098153Z digest=sha256:5fb59e8b76a46445e4cb78fd8585ad10f5de75a3a10f56945486d3c691458c37

Observation 265d6237-4a9f-4d49-b397-06d07528c0ee · outbound

This paper cites an unresolved cited work.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Unresolved cited work

Reference 2014

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:51.992417Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:51.992417Z digest=sha256:6e9fc9799045e735b43c2d2384cd637a8f83b90e6791a8cfce238eaf860ef344

Observation 3f922e3a-e195-4d6e-baec-54f102dbca45 · outbound

This paper cites an unresolved cited work.

Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation Unresolved cited work

Reference 2023

Resolution
unresolved
no resolver link, observed 2026-08-03T22:13:50.082887Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-03T22:13:50.082887Z digest=sha256:4855c73de551afcb76da4f79bbb5b4f244d6221b81f1b8033c7b9d5f51440918

Pith citing papers

Observation e33b6d2c-f801-402b-be1b-d0f5f2b55709 · inbound

Learning Structured Robot Policies from Vision-Language Models via Synthetic Neuro-Symbolic Supervision cites this paper.

Learning Structured Robot Policies from Vision-Language Models via Synthetic Neuro-Symbolic Supervision Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-01T02:03:29.524259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-13T20:11:58.423612Z digest=sha256:0f7bc8b20b937dc6d45bd241b0667ef76490d31890069b05a335a056eda425b4

Observation d2dfba28-d0b1-4370-91e8-61f472f0de65 · inbound

Learning Structured Robot Policies from Vision-Language Models via Synthetic Neuro-Symbolic Supervision cites this paper.

Learning Structured Robot Policies from Vision-Language Models via Synthetic Neuro-Symbolic Supervision Synthetic Stimuli, Real Gains: Rethinking VLM Fine-Tuning Through Fully Controlled Data Generation

Reference 11

Resolution
verified exact
arxiv_id, observed 2026-06-01T02:03:29.524259Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-08T06:32:00.761636+00:00.

source=pdf_text observed=2026-05-19T17:30:08.755545Z digest=sha256:7e819b39c714e082b7cef57ed94bcc2d49dbc1068b98fa41b51c99f5b26d21f7