Pith. sign in

Paper Citation Record · LEDGER

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images

As of 22 August 2026, this Paper Citation Record lists 52 of 52 outbound references and 0 inbound Pith citation observations for arXiv:2508.21565.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2508.21565 v1

Coverage vector

measured 52 of 52 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-15T16:43:52.244483Z

measured 52 of 52 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-22T06:32:14.747728+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

52 of 52 outbound references displayed

  • verified exact0
  • verified fuzzy40
  • unresolved12
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation fa73e3e2-8e30-475d-92a4-f9dd7715655b · outbound

This paper cites Google street view: Capturing the world at street level.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Google street view: Capturing the world at street level

Reference 1

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.663051Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:51.819293Z digest=sha256:3fca9eaac05dd07f054cdb15c9259a0511d57e0be004636955b42eaa7c46f9cd

Observation 3aab4094-9239-493d-864c-2edac5b47858 · outbound

This paper cites Evaluation methods for landscapes with greenery.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Evaluation methods for landscapes with greenery

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.639234Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:51.827290Z digest=sha256:b4b77ffde9dba42c915d1b9d53e94ace5280d568287a157baf583355aa1b5eba

Observation a1fb3897-2151-4dca-947c-80252575d21d · outbound

This paper cites Browning, Jiaying Dong, Kuiran Zhang, Shuai Yuan, H¨useyin Ertan ˙Inan, Olivia McAnirlin, Dani T.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Browning, Jiaying Dong, Kuiran Zhang, Shuai Yuan, H¨useyin Ertan ˙Inan, Olivia McAnirlin, Dani T

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.612154Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:51.847823Z digest=sha256:6c17e23192c597231f38e23cccc4aaa403d030cee950153451f79559f6e4d528

Observation e8d66e39-f41c-437c-a787-4cba57e8efd7 · outbound

This paper cites The green window view index: automated multi-source visi- bility analysis for a multi-scale assessment of green window views.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images The green window view index: automated multi-source visi- bility analysis for a multi-scale assessment of green window views

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.589324Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:51.859841Z digest=sha256:008c2690c7b0e6ea871c5ac7a4b2c4c357c526a0b9b810be404530e7dc207a0d

Observation 74d718cc-de8b-481c-9e96-cbea460508ac · outbound

This paper cites Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi- ancarlo Baldan, and Oscar Beijbom.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Gi- ancarlo Baldan, and Oscar Beijbom

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:51.870945Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:51.870945Z digest=sha256:978104250fa6d6679cba240392b914dfc49e088b691817734b873c0d5b5afb13

Observation afb504ce-f976-4375-bcc5-24f1e60f4a60 · outbound

This paper cites End-to- end object detection with transformers.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images End-to- end object detection with transformers

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.546217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:51.879017Z digest=sha256:004199b0e59d3aa729812c9eb60eadb3946a484e550d08a02ad105704e99c049

Observation c147cb24-a1c3-4ea8-bd4f-47b197d61d61 · outbound

This paper cites Spatialvlm: Endowing vision-language models with spatial reasoning capabilities.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Spatialvlm: Endowing vision-language models with spatial reasoning capabilities

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.521387Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:51.886291Z digest=sha256:dabe00e76a776e8daafbb9761b0242f1775d8d34ffca4ab0407d2df5b63d7f55

Observation a1ea861b-0d50-4ccb-a644-ed1ec0c13e3a · outbound

This paper cites Evaluating implied urban nature vital- ity in san francisco: An interdisciplinary approach combining census data, street view images, and social media analysis.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Evaluating implied urban nature vital- ity in san francisco: An interdisciplinary approach combining census data, street view images, and social media analysis

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.500260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:51.893409Z digest=sha256:e1ee28031fed2f0789c69dc836cc71cfeda6a69e74bcba94bea94c4404fbf24a

Observation 9251068b-ef49-40f2-95f1-c7f4d905fe27 · outbound

This paper cites Visual chain- of-thought prompting for knowledge-based visual reasoning.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Visual chain- of-thought prompting for knowledge-based visual reasoning

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.479241Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:51.899935Z digest=sha256:ea940dc2ac4987be2a6cae6779a204c647e2a282edb3dc8a1b96b3c66da38329

Observation 034e278e-a75c-4733-bb7b-0bc140cb65fd · outbound

This paper cites The cityscapes dataset for semantic urban scene understanding.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images The cityscapes dataset for semantic urban scene understanding

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.450500Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:51.907753Z digest=sha256:af46add3184e2568a535d98fc016ebee1d575a611ce9dec6a89d42b87ffe5990

Observation 8ac9be47-7346-48d5-963b-985d87dbaeb4 · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Instructblip: Towards general-purpose vision-language models with instruction tuning

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.421217Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:51.916567Z digest=sha256:6f68b552c7e700edfed8665cdca9401ef0cf496f6d769f469fa7432ae52afb1a

Observation 30e2affe-ff1b-4eac-af09-77c67577f38a · outbound

This paper cites Feeling Nature: Measuring perceptions of biophilia across global biomes using visual AI,.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Feeling Nature: Measuring perceptions of biophilia across global biomes using visual AI,

Reference 12

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.401105Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:51.922767Z digest=sha256:4979b754b42cf148952dabe03438776f80dac8c92c8d6fdc4cee2d2ae14b1656

Observation 59eb0f1b-c05a-4a86-b9c0-572a2d45cd4d · outbound

This paper cites an unresolved cited work.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Unresolved cited work

Reference 13

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:43:53.380042Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:51.931046Z digest=sha256:4a55eb49d8e5ade6fbb1f23dab7cd9a9f7aa98a63fd5f3c7e7281f24008f0219

Observation 896a4ee2-26e0-48dc-a808-025e6a3038c5 · outbound

This paper cites The processing of negation and polarity: An overview.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images The processing of negation and polarity: An overview

Reference 14

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.339939Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:51.938652Z digest=sha256:4227cec8e472cbab6d0bacaec62fda62cbea604324c43791fb24769a13a143ed

Observation 1e2386f4-122c-4e00-96f1-806d80b07ccd · outbound

This paper cites Vqa-lol: Visual question answering under the lens of logic.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Vqa-lol: Visual question answering under the lens of logic

Reference 15

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.318735Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:51.952079Z digest=sha256:736f10db532e4a732551e1e26fa0328d9804066d7c1a00ceab7345afaf7b6bb5

Observation e4a42176-215c-4c1e-9e75-0355d00db18d · outbound

This paper cites an unresolved cited work.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Unresolved cited work

Reference 16

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:43:53.299469Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:51.963718Z digest=sha256:f2c121b2e91725fa4ec4287740ebc051bec15025fe205fbab0cda23eb086f169

Observation a27f872f-ccec-46c1-8cf2-84ec2b934bd4 · outbound

This paper cites an unresolved cited work.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Unresolved cited work

Reference 17

Resolution
unresolved
raw_fallback, observed 2026-08-15T16:43:53.278062Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:51.969332Z digest=sha256:32a6de132d8c1e62d94180c342f94c23b40ccdf3598900c85472e2d42fa18459

Observation 8f65b9f8-92b4-47b9-a2ed-7820c2fceb60 · outbound

This paper cites Natural Adversarial Examples.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Natural Adversarial Examples

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:51.975849Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:51.975849Z digest=sha256:dd9f66f1667a66880cdc822ea101da2530721e5eb431daed8b944feb0f37b593

Observation ad873a86-b2d6-4301-afae-49cd838a606f · outbound

This paper cites Which street is hotter? street morphol- ogy may hold clues -thermal environment mapping based on street view imagery.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Which street is hotter? street morphol- ogy may hold clues -thermal environment mapping based on street view imagery

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.250774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:51.984863Z digest=sha256:123beddcfef02f639d66c4100f5cf5ab43004131b071968411b33d32af2394f5

Observation c20955f7-af37-40ca-b428-c52899b1d746 · outbound

This paper cites Hudson and Christopher D.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Hudson and Christopher D

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.222186Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:51.993458Z digest=sha256:6c020a6a9787c0cd8a992aa431b4a8d3db424affa7f1045b8ae192a918e8e924

Observation ef39f770-dd50-4905-b031-0f75a8d8f645 · outbound

This paper cites Lawrence Zitnick, and Ross Girshick.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Lawrence Zitnick, and Ross Girshick

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.194527Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.001857Z digest=sha256:e2322633e3190097a6c98b185834fbb3b5050535780821dd8b4562691f3e5d4d

Observation a516bb51-ec3e-4461-9433-bf4296dd86a9 · outbound

This paper cites Negated and Misprimed Probes for Pretrained Language Models: Birds Can Talk, But Cannot Fly.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Negated and Misprimed Probes for Pretrained Language Models: Birds Can Talk, But Cannot Fly

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:52.008408Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:52.008408Z digest=sha256:e2ad42153ed5fdd0bf421ce804bbf84dbc84c2d31e7955a6a539daf8f12af9e6

Observation f1b6b3ae-0fe0-4dbc-a7a4-73c76d0983dc · outbound

This paper cites Large language models are zero-shot reasoners.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Large language models are zero-shot reasoners

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.162260Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.015411Z digest=sha256:0a03c8010f5dc1d4a61ae25b11fe99964125b740703510f416b4bb6947309366

Observation 321bfdfc-88b5-43fd-9f77-bad28e8087ba · outbound

This paper cites Understanding counterfactuality: A review of experimental evidence for the dual meaning of counterfactuals.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Understanding counterfactuality: A review of experimental evidence for the dual meaning of counterfactuals

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.135549Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.021415Z digest=sha256:5173fc9d5b87a1c06e878103d32a893980104264db2145b89e482933239071e9

Observation 73401931-cbf7-4d71-ab53-b03723f9f749 · outbound

This paper cites Blip- 2: Bootstrapping language-image pre-training with frozen im- age encoders and large language models.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Blip- 2: Bootstrapping language-image pre-training with frozen im- age encoders and large language models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.111299Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.027458Z digest=sha256:cfee9b3ffb0588cc70eec6e7e56e757de8601fb82120748554f6d986145c358d

Observation 8a271aaf-a8e8-4376-8eda-b2f5c4d56c29 · outbound

This paper cites Assessing street-level urban greenery using google street view and a modified green view index.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Assessing street-level urban greenery using google street view and a modified green view index

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.078958Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.034515Z digest=sha256:908c6f4a020e78105d6486d737f6318c8400c6b2e61a7e6c5e05f25912d8e27e

Observation 73a49552-63b5-4c00-9117-49fdd9df3423 · outbound

This paper cites Generalizing vision-language models to novel domains: A comprehensive survey.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Generalizing vision-language models to novel domains: A comprehensive survey

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:52.041500Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:52.041500Z digest=sha256:07719a2eea3cb9530678832f09ccda9f9ed92c63f1d8c8dfd340c6560e8ce2e7

Observation 57ba5a6c-0101-40fd-bde1-f45f0b336c66 · outbound

This paper cites Eyes can deceive: Benchmarking counterfactual rea- soning abilities of multi-modal large language models.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Eyes can deceive: Benchmarking counterfactual rea- soning abilities of multi-modal large language models

Reference 28

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.051160Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.048824Z digest=sha256:ed41c31f1b936b6f7b697a84eb9de8c8e0a7d0b0bdd54b9c3bd651a4294fc73b

Observation 3ce008d1-5d7e-4e8a-8b80-0ffb16569066 · outbound

This paper cites Evaluating human perception of building exteriors using street view imagery.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Evaluating human perception of building exteriors using street view imagery

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:53.020377Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.060380Z digest=sha256:3f728f2faec71b47c33e4315c954387dd8a0c1b5e1fdcc190deb354e8a2dbb0f

Observation 54854b56-3587-41bf-9ac8-c81ee1c67148 · outbound

This paper cites Improved baselines with visual instruction tuning.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Improved baselines with visual instruction tuning

Reference 30

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:52.066628Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:52.066628Z digest=sha256:e651651960aa2aab4ac9f9fc99e2f560b82e8fe0fc3484b06c9052fadc6a8a30

Observation cfd2e46e-acbe-4658-83bd-ff45a21775ac · outbound

This paper cites Efficacy of Synthetic Data as a Benchmark.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Efficacy of Synthetic Data as a Benchmark

Reference 31

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:52.076558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:52.076558Z digest=sha256:3f053d7243ff3862f700421980fa21e77143596216c5bc6bbbe39b6b9438cb56

Observation 5f6f332f-f551-458e-bf6d-e67a7ebae481 · outbound

This paper cites Review of methods used to es- timate the sky view factor in urban street canyons.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Review of methods used to es- timate the sky view factor in urban street canyons

Reference 32

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.972412Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.084073Z digest=sha256:b8a75efcc20c2bffe98134709778a39288abcd58f547031ef98880c3dc77f290

Observation 4d6efcc9-874f-485c-b432-09bc7612846a · outbound

This paper cites Objective scoring of streetscape walkability related to leisure walking: Statistical modeling approach with semantic segmentation of google street view images.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Objective scoring of streetscape walkability related to leisure walking: Statistical modeling approach with semantic segmentation of google street view images

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.948854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.090882Z digest=sha256:0d2e83b9b6d06f2cbf06f196dad05aa1b0e1849df3204e751ba40adafc3740bf

Observation 8d40bd89-9d58-421e-bacc-c3b29d53fe68 · outbound

This paper cites Streetscore - predicting the perceived safety of one mil- lion streetscapes.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Streetscore - predicting the perceived safety of one mil- lion streetscapes

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.920991Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.096672Z digest=sha256:27f17faf11f252f12f9cf7d44f1e8dd9f37d8176ca5a9f63dedacb4d689491ea

Observation f7984d46-b145-4026-bdba-89620c64ac7c · outbound

This paper cites The mapillary vistas dataset for semantic understanding of street scenes.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images The mapillary vistas dataset for semantic understanding of street scenes

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.897550Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.102021Z digest=sha256:d60a88ce7fc6b15e6f795f3008caa796b08adee7353ace0c590a38a40568422e

Observation 714a7014-80c6-4efc-8125-cbfcf2e43b1b · outbound

This paper cites Counterfactual vqa: A cause- effect look at language bias.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Counterfactual vqa: A cause- effect look at language bias

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.875949Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.108525Z digest=sha256:d3cf7b73383dd23f5f1efab123fa9e5d26d78d3304bf638df39e8fce5a36cf3f

Observation 565549e6-5b89-4887-9a93-0608d7d0247d · outbound

This paper cites Evaluating the subjective percep- tions of streetscapes using street-view images.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Evaluating the subjective percep- tions of streetscapes using street-view images

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.840854Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.117459Z digest=sha256:484e096bf62be106b0775c80a21e40db0f3638f1f18472d28f812365652eadc2

Observation 5c478263-af87-4d41-9a62-0384e917d7c4 · outbound

This paper cites Learning transferable visual models from natural language supervision.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Learning transferable visual models from natural language supervision

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.810090Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.124226Z digest=sha256:8d0796efe8bfcd3b4fa5fd3eaaaa5cc175991ca1d02c19c72c652d922670c32c

Observation f4217137-e637-49d5-baad-d2f8a655fbf9 · outbound

This paper cites Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Towards robust monocular depth estimation: Mixing datasets for zero-shot cross-dataset transfer

Reference 39

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.783106Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.134226Z digest=sha256:4e94bcfbc8e388de49f15b7c58b556d1092dbf9b36998a9092841e4a24c4d51d

Observation c7917aaf-da19-4f57-91ae-3ff0fdab65e3 · outbound

This paper cites Visual cot: Advancing multi-modal language models with a compre- hensive dataset and benchmark for chain-of-thought reason- ing.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Visual cot: Advancing multi-modal language models with a compre- hensive dataset and benchmark for chain-of-thought reason- ing

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.749086Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.145483Z digest=sha256:213ced49e5904258347627a6a69cdce6a463b0fb17924e46d5f08bc37a7a6d55

Observation d06398fc-d73b-4ba4-aaa4-61ac3c064d4d · outbound

This paper cites Is synthetic data all we need? benchmarking the robustness of models trained with synthetic images.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Is synthetic data all we need? benchmarking the robustness of models trained with synthetic images

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.724551Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.151861Z digest=sha256:1e52594a66fb289243299a5141532ac42042ec72bd018a1c27991d2ceb9f9119

Observation 314146b7-e38e-4fac-9014-88c5c09a5ab6 · outbound

This paper cites Svensson.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Svensson

Reference 42

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.698831Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.157932Z digest=sha256:fadd507d514141e15a36f27473c08ebf6092c3fd0551f8a36648e91fc78c28c1

Observation c045f472-8ee5-47ff-8d08-9841f4de7986 · outbound

This paper cites A new benchmark: On the util- ity of synthetic data with blender for bare supervised learn- ing and downstream domain adaptation.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images A new benchmark: On the util- ity of synthetic data with blender for bare supervised learn- ing and downstream domain adaptation

Reference 43

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.663049Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.165500Z digest=sha256:1f4ad578cbed47cbdc106f1a5b564da51011e14e01d1b4db3114eb96fd64954f

Observation 5ed35ed8-5924-49c6-82f0-ecdeb80ea2b2 · outbound

This paper cites Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Gemini 1.5: Unlocking multimodal understanding across millions of tokens of context

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:52.173641Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:52.173641Z digest=sha256:e6af09e701ddeca0fd2d422a88cb642f0f5d769a37fc05a332015fe4875bf35a

Observation 31c188cc-dbb3-4a3c-b4c1-5d5c0a230415 · outbound

This paper cites Negation: A Pink Elephant in the Large Language Models' Room?.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Negation: A Pink Elephant in the Large Language Models' Room?

Reference 45

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:52.181805Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:52.181805Z digest=sha256:fee43bc29afe84717d315bdfb86314a1b0ef7698d433256d137c57adfb30fac9

Observation d676c83d-a650-4194-9a1e-d672822df9e4 · outbound

This paper cites Al- varez.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Al- varez

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.630835Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.187965Z digest=sha256:2fbb312b00ee583b60dcb89e38cc5329df93ed4b4191135e4cd15d08f76413cb

Observation a065ad02-441c-4927-83fa-cdfbcdeeb656 · outbound

This paper cites Self-consistency improves chain of thought reasoning in lan- guage models.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Self-consistency improves chain of thought reasoning in lan- guage models

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.598326Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.195234Z digest=sha256:4f748890d29e27ad2c7b5fdc217ea970f7adf89e36fc541a78cd494d11f5d8aa

Observation ff0f101b-8ba0-49a6-833d-513bb6a8efaa · outbound

This paper cites Chi, Quoc V.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Chi, Quoc V

Reference 48

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.572855Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.201830Z digest=sha256:673475b31e9c14f44caf908f244a326be2253a38412e156bd5d1fba3e23165c8

Observation 0802696a-1af0-4e2f-8a0a-665d74fa90cc · outbound

This paper cites Alvarez, and Ping Luo.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Alvarez, and Ping Luo

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.536709Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.212828Z digest=sha256:892d30d726d3fab6e4ac2153a763d4f80038eb127a5e765c86e17cadaa627ae9

Observation 20365222-5643-4cc0-89c8-6a25d8ace7c8 · outbound

This paper cites Improve vision language model chain-of-thought reasoning, 2024.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images Improve vision language model chain-of-thought reasoning, 2024

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.514988Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.219179Z digest=sha256:4259eb219352dbe0f9d98c69e0a0766cbbae60638f9889f33b6e048b5931d8ef

Observation df454ef8-5092-4cbf-81b2-d66f19ccb8fa · outbound

This paper cites NegVQA: Can Vision Language Models Understand Negation?.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images NegVQA: Can Vision Language Models Understand Negation?

Reference 51

Resolution
unresolved
no resolver link, observed 2026-08-15T16:43:52.235796Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-15T16:43:52.235796Z digest=sha256:29de6b75bc0d3839db434827ab664345f09fa43435d2b7729f6afbae03c4d34a

Observation f516c9c9-7109-428c-9898-a64ab16fc56e · outbound

This paper cites A study on the impact of visible green index and vegetation structures on brain wave change in residential landscape.Urban Forestry & Urban Greening, 64:127299, 2021.

How Well Do Vision--Language Models Understand Cities? A Comparative Study on Spatial Reasoning from Street-View Images A study on the impact of visible green index and vegetation structures on brain wave change in residential landscape.Urban Forestry & Urban Greening, 64:127299, 2021

Reference 52

Resolution
verified fuzzy
raw_fallback, observed 2026-08-15T16:43:52.491168Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-22T06:32:14.747728+00:00.

source=pdf_text observed=2026-08-15T16:43:52.244483Z digest=sha256:8d17bd65ea2451d970674f67c08b614d15edc2f31ff1f0b7cd78c1b92d27bd33

Pith citing papers

No inbound Pith citation observations are available.