Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T16:43:52.244483Z
Paper Citation Record · LEDGER
As of 18 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2508.21565.
A citation records a reference. It does not transfer a finding from one paper to another.
Typed states for the displayed outbound observations.
Source: paper_references, paper_reference_links, observed 2026-08-15T16:43:52.244483Z
One-hop event checks from named stored sources.
Source: scholarly_work_events, retraction_status_cache, observed 2026-08-17T06:30:58.91139+00:00
Pith citing papers itemized under the disclosed page cap.
Source: paper_references, paper_reference_links
A source-named dated measurement, never combined with another source.
Source: cited_works
52 of 52 outbound references displayed
External citation measurements
No source-named external measurement is stored.
Observation fa73e3e2-8e30-475d-92a4-f9dd7715655b · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Google street view: Capturing the world at street level
Reference 1
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3aab4094-9239-493d-864c-2edac5b47858 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Evaluation methods for landscapes with greenery
Reference 2
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a1fb3897-2151-4dca-947c-80252575d21d · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Browning, Jiaying Dong, Kuiran Zhang, Shuai Yuan, H¨useyin Ertan ˙Inan, Olivia McAnirlin, Dani T
Reference 3
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e8d66e39-f41c-437c-a787-4cba57e8efd7 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images The green window view index: automated multi-source visi- bility analysis for a multi-scale assessment of green window views
Reference 4
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 74d718cc-de8b-481c-9e96-cbea460508ac · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi- ancarlo Baldan, and Oscar Beijbom
Reference 5
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation afb504ce-f976-4375-bcc5-24f1e60f4a60 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images End-to- end object detection with transformers
Reference 6
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c147cb24-a1c3-4ea8-bd4f-47b197d61d61 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Spatialvlm: Endowing vision-language models with spatial reasoning capabilities
Reference 7
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a1ea861b-0d50-4ccb-a644-ed1ec0c13e3a · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Evaluating implied urban nature vital- ity in san francisco: An interdisciplinary approach combining census data, street view images, and social media analysis
Reference 8
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 9251068b-ef49-40f2-95f1-c7f4d905fe27 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Visual chain- of-thought prompting for knowledge-based visual reasoning
Reference 9
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 034e278e-a75c-4733-bb7b-0bc140cb65fd · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images The cityscapes dataset for semantic urban scene understanding
Reference 10
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8ac9be47-7346-48d5-963b-985d87dbaeb4 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Instructblip: Towards general-purpose vision-language models with instruction tuning
Reference 11
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 30e2affe-ff1b-4eac-af09-77c67577f38a · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Feeling Nature: Measuring perceptions of biophilia across global biomes using visual AI,
Reference 12
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 59eb0f1b-c05a-4a86-b9c0-572a2d45cd4d · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Unresolved cited work
Reference 13
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 896a4ee2-26e0-48dc-a808-025e6a3038c5 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images The processing of negation and polarity: An overview
Reference 14
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 1e2386f4-122c-4e00-96f1-806d80b07ccd · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Vqa-lol: Visual question answering under the lens of logic
Reference 15
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation e4a42176-215c-4c1e-9e75-0355d00db18d · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Unresolved cited work
Reference 16
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a27f872f-ccec-46c1-8cf2-84ec2b934bd4 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Unresolved cited work
Reference 17
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8f65b9f8-92b4-47b9-a2ed-7820c2fceb60 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Natural Adversarial Examples
Reference 18
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation ad873a86-b2d6-4301-afae-49cd838a606f · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Which street is hotter? street morphol- ogy may hold clues -thermal environment mapping based on street view imagery
Reference 19
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c20955f7-af37-40ca-b428-c52899b1d746 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Hudson and Christopher D
Reference 20
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ef39f770-dd50-4905-b031-0f75a8d8f645 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Lawrence Zitnick, and Ross Girshick
Reference 21
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a516bb51-ec3e-4461-9433-bf4296dd86a9 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Negated and Misprimed Probes for Pretrained Language Models: Birds Can Talk, But Cannot Fly
Reference 22
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f1b6b3ae-0fe0-4dbc-a7a4-73c76d0983dc · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Large language models are zero-shot reasoners
Reference 23
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 321bfdfc-88b5-43fd-9f77-bad28e8087ba · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Understanding counterfactuality: A review of experimental evidence for the dual meaning of counterfactuals
Reference 24
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 73401931-cbf7-4d71-ab53-b03723f9f749 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Blip- 2: Bootstrapping language-image pre-training with frozen im- age encoders and large language models
Reference 25
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8a271aaf-a8e8-4376-8eda-b2f5c4d56c29 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Assessing street-level urban greenery using google street view and a modified green view index
Reference 26
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 73a49552-63b5-4c00-9117-49fdd9df3423 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Generalizing vision-language models to novel domains: A comprehensive survey
Reference 27
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 57ba5a6c-0101-40fd-bde1-f45f0b336c66 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Eyes can deceive: Benchmarking counterfactual rea- soning abilities of multi-modal large language models
Reference 28
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 3ce008d1-5d7e-4e8a-8b80-0ffb16569066 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Evaluating human perception of building exteriors using street view imagery
Reference 29
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 54854b56-3587-41bf-9ac8-c81ee1c67148 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Improved baselines with visual instruction tuning
Reference 30
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation cfd2e46e-acbe-4658-83bd-ff45a21775ac · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Efficacy of Synthetic Data as a Benchmark
Reference 31
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 5f6f332f-f551-458e-bf6d-e67a7ebae481 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Review of methods used to es- timate the sky view factor in urban street canyons
Reference 32
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 4d6efcc9-874f-485c-b432-09bc7612846a · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Objective scoring of streetscape walkability related to leisure walking: Statistical modeling approach with semantic segmentation of google street view images
Reference 33
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 8d40bd89-9d58-421e-bacc-c3b29d53fe68 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Streetscore - predicting the perceived safety of one mil- lion streetscapes
Reference 34
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f7984d46-b145-4026-bdba-89620c64ac7c · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images The mapillary vistas dataset for semantic understanding of street scenes
Reference 35
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 714a7014-80c6-4efc-8125-cbfcf2e43b1b · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Counterfactual vqa: A cause- effect look at language bias
Reference 36
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 565549e6-5b89-4887-9a93-0608d7d0247d · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Evaluating the subjective percep- tions of streetscapes using street-view images
Reference 37
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5c478263-af87-4d41-9a62-0384e917d7c4 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Learning transferable visual models from natural language supervision
Reference 38
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation f4217137-e637-49d5-baad-d2f8a655fbf9 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer
Reference 39
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c7917aaf-da19-4f57-91ae-3ff0fdab65e3 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Visual cot: Advancing multi-modal language models with a compre- hensive dataset and benchmark for chain-of-thought reason- ing
Reference 40
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation d06398fc-d73b-4ba4-aaa4-61ac3c064d4d · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Is synthetic data all we need? benchmarking the robustness of models trained with synthetic images
Reference 41
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 314146b7-e38e-4fac-9014-88c5c09a5ab6 · outbound
Reference 42
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation c045f472-8ee5-47ff-8d08-9841f4de7986 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images A new benchmark: On the util- ity of synthetic data with blender for bare supervised learn- ing and downstream domain adaptation
Reference 43
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 5ed35ed8-5924-49c6-82f0-ecdeb80ea2b2 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context
Reference 44
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation 31c188cc-dbb3-4a3c-b4c1-5d5c0a230415 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Negation: A Pink Elephant in the Large Language Models' Room?
Reference 45
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation d676c83d-a650-4194-9a1e-d672822df9e4 · outbound
Reference 46
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation a065ad02-441c-4927-83fa-cdfbcdeeb656 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Self-consistency improves chain of thought reasoning in lan- guage models
Reference 47
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation ff0f101b-8ba0-49a6-833d-513bb6a8efaa · outbound
Reference 48
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 0802696a-1af0-4e2f-8a0a-665d74fa90cc · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Alvarez, and Ping Luo
Reference 49
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation 20365222-5643-4cc0-89c8-6a25d8ace7c8 · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Improve vision language model chain-of-thought reasoning, 2024
Reference 50
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
Observation df454ef8-5092-4cbf-81b2-d66f19ccb8fa · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images NegVQA: Can Vision Language Models Understand Negation?
Reference 51
Source-reported events for the cited work
Unavailable: canonical work link unavailable.
Observation f516c9c9-7109-428c-9898-a64ab16fc56e · outbound
How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images A study on the impact of visible green index and vegetation structures on brain wave change in residential landscape.Urban Forestry & Urban Greening, 64:127299, 2021
Reference 52
Source-reported events for the cited work
No event found in the named queried sources as of 2026-08-17T06:30:58.91139+00:00.
No inbound Pith citation observations are available.