Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:14:42.975394Z
Paper Citation Record · LEDGER
As of 19 August 2026, this Paper Citation Record lists 34 of 34 outbound references and 1 inbound Pith citation observation for arXiv:2505.10604.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T21:14:42.975394Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links, observed 2026-07-15T13:03:40.459461Z
A source-named dated measurement, never combined with another source.
Source: cited_works
34 of 34 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation 87e88345-65e0-419a-b17c-5659c3503a55 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Flamingo: a visual language model for few-shot learning, 2022
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 07686218-6953-49cc-840b-de759a06fcc6 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Qwen-vl: A versatile vision-language model for understanding, localization, text reading, and beyond, 2023
Reference 2
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 6ad4ae12-e25b-4d13-a6c5-086cad8308df · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Spatialbot: Precise spatial understanding with vision language models, 2025
Reference 3
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 550d2283-d28e-4a7d-b7c5-40d966de3523 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence How far are we to gpt-4v? closing the gap to commercial multimodal models with open-source suites, 2024
Reference 4
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 996e1359-3b3e-4eb9-8385-dd6c5f6d13a4 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Internvl: Scaling up vision foundation models and aligning for generic visual-linguistic tasks, 2024
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation be14955a-a184-4eaa-b759-de79b3014db8 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Scaling egocentric vision: The epic-kitchens dataset
Reference 6
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation c046bf85-b54c-416a-ba39-7949a06a74bd · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Geobench-vlm: Benchmarking vision-language models for geospatial tasks, 2025
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 17266c76-1d65-4af0-98af-028779f3b487 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Mm-spatial: Exploring 3d spatial understanding in multimodal llms, 2025
Reference 8
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 42e69c59-eac7-4258-8f7d-b19164ff9dde · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Patch n’ pack: Navit, a vision transformer for any aspect ratio and resolution, 2023
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation e0e633db-67bb-4eca-bc9a-d3e91b5be9c2 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Smith, Hannaneh Hajishirzi, Ross Girshick, Ali Farhadi, and Aniruddha Kembhavi
Reference 10
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31bfc3c0-47aa-4d7c-ba77-88b5c6288cfd · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence An image is worth 16x16 words: Transformers for image recognition at scale, 2021
Reference 11
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation e8723a1f-e289-433b-bbe3-ce8d23587007 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Minicpm: Unveiling the potential of small language models with scalable training strategies, 2024
Reference 12
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ea6c64e2-2557-452d-b62e-c97c500880d9 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Blip: Bootstrapping language-image pre-training for unified vision-language understanding and generation, 2022
Reference 13
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation bf4a7934-a9fb-45f8-b5ae-e428bc15bf20 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Sti- bench: Are mllms ready for precise spatial-temporal world understanding?, 2025
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a5127df8-e170-47cd-8905-165b1a3d83cc · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Visual instruction tuning, 2023
Reference 15
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 1499ebe9-a050-4dd8-819a-34749f8621e0 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence ivispar – an interactive visual-spatial reasoning benchmark for vlms, 2025
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6a684c88-8b16-42f6-8533-24566499da66 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Learning transferable visual models from natural language supervision, 2021
Reference 17
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d21812cd-ad75-485a-8782-3a566f5efc54 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Gsr-bench: A benchmark for grounded spatial reasoning evaluation via multimodal llms, 2024
Reference 18
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d5289e28-15f5-4bcf-abf0-e7dd2e3a4887 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Unresolved cited work
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 2f9c5f1d-a126-44e2-a7a3-a61d24c260c1 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Qwen2-vl: Enhancing vision- language model’s perception of the world at any resolution, 2024
Reference 20
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 96bb1d28-a914-49dc-8036-b391f3124010 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence ST-Think: How Multimodal Large Language Models Reason About 4D Worlds from Ego-Centric Videos
Reference 21
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 13f3de2b-29b2-439c-b235-d8ccd82f46b4 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Lvlm-ehub: A comprehensive evaluation benchmark for large vision-language models, 2023
Reference 22
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 508ec8fc-6d1a-449d-811f-9db2594959f3 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Gupta, Rilyn Han, Li Fei-Fei, and Saining Xie
Reference 23
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 56015db9-22e8-4f78-868c-05ac3d0a976f · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Minicpm-v: A gpt-4v level mllm on your phone, 2024
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 7a48a703-c9b7-47da-ac1f-d364b701f184 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Good at captioning, bad at counting: Benchmarking gpt-4v on earth observation data, 2024
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7869ba2d-d9e5-4933-a60e-25f151893c0c · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence image_caption
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation de0e1514-0314-414e-9cd3-c2ddbb478ab8 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Unresolved cited work
Reference 27
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 6cf4dfac-5101-4e18-940b-1c504da11ffc · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Unresolved cited work
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 10f6f34c-a698-4d3c-b115-630d2ef5a125 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Unresolved cited work
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation d18fda69-9fef-406e-9282-80c02c983538 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Unresolved cited work
Reference 30
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 602e3c84-08a4-4704-a2ad-0a28f95d1a03 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence } Stage 2: Task-Specific Questions a. Spatial Relation Task RELATION BASE PROMPT You should output a json string with format {
Reference 31
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 663d2553-90c8-4201-a73d-d98f83ccf89f · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence 3Implemented with PIL.Image.transpose
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 64ddf566-fd70-4027-8179-40acbeb1c2d9 · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Limitations
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation 7c749a1e-b4b1-4bbb-a9ce-a288f0d04e9e · outbound
MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence Guidelines: • The answer NA means that the paper does not involve crowdsourcing nor research with human subjects
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.
Observation a09fba45-765b-4fce-a3cf-5210b421166d · inbound
It's Time to Get It Right: Improving Analog Clock Reading and Clock-Hand Spatial Reasoning in Vision-Language Models MIRAGE: A Multi-modal Benchmark for Spatial Perception, Reasoning, and Intelligence
Reference 24
Source-reported events for the cited work
Unavailable: canonical work link unavailable.