Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:50:34.081389Z
Paper Citation Record · LEDGER
As of 7 August 2026, this Paper Citation Record lists 42 of 42 outbound references and 1 inbound Pith citation observation for arXiv:2505.20728.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-07T13:50:34.081389Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-07T06:34:17.273281+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-05-08T06:34:56.032634Z
A source-named dated measurement, never combined with another source.
Source: arxiv_reference, observed 2026-05-11T21:11:15.085382Z
42 of 42 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 5c399578-31b0-470e-9f41-1b4bc570dda5 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models online" 'onlinestring :=
Reference 1
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 9d151fad-124c-448c-802b-279daa1da518 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models write newline
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation b7040f26-b49f-46db-9f98-0ba1569ab0bb · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Phi-4-Mini Technical Report: Compact yet Powerful Multimodal Language Models via Mixture-of-LoRAs
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 752de2f1-8c9d-4fd0-b9d1-e72e290aab7a · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models GPT-4 Technical Report
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 915d26d0-19a8-4d7f-8af0-09413637133e · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Gemini: A Family of Highly Capable Multimodal Models
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation a998c525-c809-401b-a93c-920b0ea44347 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5d9339a7-f16c-43db-b718-064548533b5c · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Qwen2.5-VL Technical Report
Reference 7
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 18f9b066-a05a-4a05-8efe-ba769070c060 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 83229649-630f-4331-9b98-ad31b15475fc · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0b6d4f18-972e-4d8c-b198-074d5dd7cb52 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Aya Vision: Advancing the Frontier of Multilingual Multimodality
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 88c75eb4-923e-4bb7-bbb5-8cb2c3f30022 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Kimi-VL Technical Report
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d67f9f9f-cb3a-43ab-87b8-be3144fcc46f · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 4b860dae-a052-4a8b-bf93-b134355876e7 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models MME: A Comprehensive Evaluation Benchmark for Multimodal Large Language Models
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1e2bdc44-4582-4ee6-866a-cbbed726afb1 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 56b1f494-4cca-4dd8-8989-6add8b57b1e8 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 06300119-8468-4e4b-b66c-f7cbd0fe8a1f · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work
Reference 16
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8d8cc75a-6a49-4d13-b764-86b738695d19 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1751a488-47f6-463a-95ee-3434fdd5a170 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 98096c57-9bd8-444f-a72b-8ac62487f656 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work
Reference 19
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 8835969b-b7aa-46d7-ab34-527f0d1550ba · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e95c8114-bbb1-485f-a17f-6ce584870e4e · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 001d6cb3-83f1-49ba-98c9-d03602fda910 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 601b543a-3e8e-40a9-ba87-49474f2fca34 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 0f972e68-e574-42ee-a6ed-21d647e6ff52 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models CAPTURe: Evaluating Spatial Reasoning in Vision Language Models via Occluded Object Counting
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ead65715-e649-4dca-b20c-e1f20b6c4f0f · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5c97cc13-0618-4740-880f-af2c42d94a96 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work
Reference 26
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ffbe223e-d9c3-4ae9-9117-c5fa30c22be1 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models VGRP-Bench: Visual Grid Reasoning Puzzle Benchmark for Large Vision-Language Models
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dfb006dd-2a19-4373-a98d-e6b997626f22 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models LAION-400M: Open Dataset of CLIP-Filtered 400 Million Image-Text Pairs
Reference 28
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 874f6b88-88ce-4ec7-a840-424a733f8385 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work
Reference 29
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 4eae3f36-1e9a-4657-b9a4-c3b516b3d359 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 62794210-eb1e-45bc-b7fd-b271d7ac7dc9 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation dc85e25c-5131-46e0-a594-a8fb0bacfe77 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models VisualPuzzles: Decoupling Multimodal Reasoning Evaluation from Domain Knowledge
Reference 32
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation fc5d705f-b0e8-4de9-a506-d73312103c29 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Mind the Gap: Benchmarking Spatial Reasoning in Vision-Language Models
Reference 33
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f462880d-35a6-456c-a477-9290ac0be834 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models LEGO-Puzzles: How Good Are MLLMs at Multi-Step Spatial Reasoning?
Reference 34
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d5a87d20-629c-4fed-a2ca-1aa7cc8a503d · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Crossmodal-3600: A Massively Multilingual Multimodal Evaluation Dataset
Reference 35
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cf46d1fe-c7ad-4a60-ae0d-05f05bc2541a · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 5594654b-632a-47ac-b866-0c18d131b687 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Qwen2-VL: Enhancing Vision-Language Model's Perception of the World at Any Resolution
Reference 37
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 537d74e7-7f34-4249-9156-1de218b34bdc · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models CogLM: Tracking Cognitive Development of Large Language Models
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.
Observation 1633c790-48b4-4a8f-ad8c-b011cab5840c · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work
Reference 39
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ae6ea074-f45a-4264-af88-2e2f6d623eb0 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Unresolved cited work
Reference 40
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 01d53ced-5ee7-4b5e-b343-a35ca024c4f8 · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models Redundancy Principles for MLLMs Benchmarks
Reference 41
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d9652d4f-5803-404f-bfbe-04cbc5d6017a · outbound
Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models InternVL3: Exploring Advanced Training and Test-Time Recipes for Open-Source Multimodal Models
Reference 42
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation af191167-4ede-47cd-87e2-f3dc88dfe650 · inbound
ShredBench: Evaluating the Semantic Reasoning Capabilities of Multimodal LLMs in Document Reconstruction Jigsaw-Puzzles: From Seeing to Understanding to Reasoning in Vision-Language Models
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-07T06:34:17.273281+00:00.