REVIEW 4 major objections 4 minor 32 references
Can a Large Language Model Assess Urban Design Quality? Evaluating Walkability Metrics Across Expertise Levels
T0 review · 4 major / 4 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Giving a multimodal language model formal, expert-written walkability metric definitions in its prompts makes its ratings of street-view images more consistent and less optimistic.
desk verdict A useful exploratory study of prompt-level expertise effects on MLLM walkability scoring, but the central evaluative-performance claim needs a human ground truth it currently lacks. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The carrying mechanism is a four-level expertise gradient built into the prompt: Level 1 asks for safety and attractiveness ratings with no metrics, Level 2 uses vague metric names from the literature, Level 3 uses quantified metric names, and Level 4 adds a formal description and scoring rule for each metric. The metric set itself is assembled from two walkability review literatures and structured through an ontology-based categorisation, with 21 safety metrics and 21 attractiveness metrics selected as comparison sets. The formal descriptions are the active ingredient: they convert short labels into operational scoring instructions, reducing the ambiguity that lets the model read crosswalks as fixed furniture or infer traffic calming devices from narrow roads. Statistical tests (Levene, Welch ANOVA, Games-Howell, Kruskal-Wallis) are used to show that the prompt conditions produce different distributions and that the fully described condition concentrates scores.
What would settle it
Have expert urban designers score the same street-view images with the same 21 safety and 21 attractiveness metrics, then compare their scores with the model's four prompt conditions after normalizing the different score scales; if human expert scores do not agree more with the fully described condition than with the no-metrics condition, the claim that expert-knowledge prompts improve evaluative performance fails. A second check is to run the same four prompts on a different multimodal model: if the concentration effect disappears, it is a property of the model rather than of expert-knowledge prompting.
Extended reading notes
Core claim
The paper's central claim is that a multimodal large language model's evaluative performance can be enhanced by integrating expert knowledge, and that increasing the semantic clarity of that knowledge improves consistency of the evaluative outputs. Concretely, when ChatGPT-4 is given no evaluation criteria it rates pedestrian safety and attractiveness more optimistically and diverges from metric-informed models; when given literature metrics, overall score distributions shift and stabilize; and when given quantified metric names plus formal descriptions with scoring rules, per-metric scores become more concentrated and the model stops making certain interpretive errors. The authors do not claim these expert-informed scores are objectively correct: comparison with human urban design practitioners is explicitly left to future work.
Load-bearing premise
The load-bearing premise is that higher consistency and concentration of the model's scores count as better evaluative performance, since the paper includes no human expert scoring or objective walkability benchmark to confirm that the more concentrated, expert-informed scores are actually right.
Editorial extensions
If this is right
- Automated walkability screening can be steered without fine-tuning: writing formal metric descriptions into the prompt materially changes how the model rates street-view images.
- Unaided multimodal language model ratings are optimism-prone, so any automated urban quality workflow that skips metric definitions should treat its scores cautiously.
- Formal descriptions reduce per-metric misinterpretation, making scores more concentrated and better comparable across streets within a metric.
- Quantified metric names alone do not significantly shift overall score distributions; the descriptive layer is what aligns the model with the intended criteria at the metric level.
- The approach can translate low scores on design-actionable metrics into targeted urban design interventions, demonstrated for two lower-scoring streets in the study.
Reading between the lines
- If the consistency gain is later confirmed against human expert raters, the same prompt-document technique could transfer to other perceptual urban qualities such as enclosure, imageability, or maintenance without retraining the model.
- Because the overall score distributions of the three metric-informed conditions showed no significant differences, the practical payoff of formal descriptions may lie in variance reduction and error correction rather than in shifting average scores; that distinction deserves an explicit test.
- A natural extension is to measure inter-run reliability by repeating each prompt condition several times, since the current single-run design cannot separate prompt-induced concentration from the model's sampling noise.
- With one model and one city's images, the safest reading is that expert-knowledge prompting changes this model's behaviour on this dataset; how far the effect extends across models, languages, and street networks remains open.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper asks whether a multimodal large language model (ChatGPT-4) can evaluate urban design quality from street-view images (SVIs), and whether injecting increasingly formal expert knowledge into the prompt changes the resulting scores. The authors collect 124 walkability metrics from the literature, select 21 safety and 21 attractiveness metrics, and construct four prompt conditions: C1 (no metrics, holistic 1–105 score), C2 (vague metric names, 1–5 per metric), C3 (quantified metric names, 1–5 per metric), and C4 (quantified metrics plus formal descriptions and scoring rules). They apply the four models to 42 SVIs from Singapore and analyze the score distributions with Levene’s test, Welch’s ANOVA, Games-Howell post-hoc tests, and Kruskal-Wallis tests. The main descriptive findings are that C1 produces more optimistic and dispersed scores, while C2–C4 are mutually closer in overall distribution, and that on selected metrics C4 yields more concentrated scores. Two example images illustrate cases where C4 avoids misinterpretations (e.g., not treating crosswalks as fixed furniture). The paper concludes that integrating expert knowledge enhances MLLMs’ evaluative performance and that increasing semantic clarity improves consistency.
Significance. If the descriptive findings hold, the paper provides a useful empirical demonstration that prompt structure and metric definition materially change MLLM-based street-environment scoring, and that well-specified rubrics reduce variance. The publication of the metric lists, prompts, and assessment data on Figshare is a reproducible resource for the urban-analytics community, and the use of inferential statistics rather than only point estimates is a strength. However, the central claim as worded in the Conclusion—that evaluative performance is enhanced—goes beyond what the data can show, because no human expert assessment or objective walkability benchmark is used. The manuscript itself defers practitioner comparison to future work. The significance of the paper therefore rests on the well-supported descriptive claims about score distributions and prompt sensitivity, not on the unsupported claim of improved correctness.
major comments (4)
- [Conclusion and §1] The claim that integrating expert knowledge enhances MLLMs’ evaluative performance is not supported by the evidence presented. The study measures consistency, concentration, and differences in score distributions; it does not measure correctness against any ground truth. The Conclusion explicitly lists ‘engaging urban design practitioners to evaluate the SVIs and compare their assessment with the MLLMs’ as future work, which acknowledges that the missing expert baseline is essential. Without such a reference, a more concentrated score distribution could reflect more rigid or systematically biased scoring, not better performance. The paper should either reframe the central claim as an effect on score distribution and consistency, or add a human-expert evaluation of the same SVIs to support the performance language.
- [§3.2, Table 3, and Figure 4] The comparison between Model-C1 and Models-C2/C3/C4 is confounded by the response scale: C1 uses a single holistic 1–105 score, while C2–C4 sum 21 metrics each scored 1–5 (also a 21–105 range, but with a different aggregation structure). Observed differences in means, variances, and rank order may therefore partly reflect the difference between holistic and decomposed scoring formats rather than the absence or presence of expert knowledge. To support the claim that expert knowledge changes evaluations, the authors should compare C1 against, for example, a decomposed but metric-free condition, or analyze normalized/standardized scores that make the two formats more comparable.
- [§4.1, Figures 5 and 6] The inference that Model-C4’s higher concentration reflects fewer misinterpretations is based on two anecdotal examples. The paper states that ‘the more varied and dispersed score distributions in Model-C3 and Model-C2 could be attributed to the ambiguity resulting from the lack of definitions,’ but this causal interpretation is not quantitatively tested. To make this load-bearing point credible, the authors should systematically classify the model’s reasoning across all 42 images per metric (for example, by coding each response as consistent or inconsistent with the provided description), rather than relying on two hand-picked cases.
- [§4.1 and Table 3] The statistical support for the central variance-related claim is partial. Table 3 shows that C2, C3, and C4 are not significantly different in overall safety and attractiveness distributions, yet the paper later emphasizes C4’s higher concentration. This is not necessarily contradictory because the metric-level analysis in Figure 5 is more fine-grained, but the paper should state clearly that the concentration effect is metric-specific and not a global property of Model-C4’s scores. The current presentation risks overgeneralizing the finding.
minor comments (4)
- [§3.2] The text says the prompts draw on metrics ‘outlined in Section 5,’ but the metrics are presented in Section 3.1 and Table 1; the cross-reference should be corrected.
- [Table 1] The entries ‘DiverseLandscape LandscapeDiversityIndex’ and ‘Colorfulness EnvironmentalColorDiversity’ appear in both the vague and quantified columns for Safety and Attractiveness, which is confusing and may be a formatting artifact. The table should clearly distinguish the two sets of metric names or explain that some names coincide.
- [§3.2] The paper states that models were tested ‘in order from Model-C1 to Model-C4’ to prevent learning from descriptions, but it is not clear whether each test used a fresh session or whether the same conversation history was retained. If a single session was used, earlier prompts could still influence later responses. Please clarify the session and conversation-reset protocol.
- [§2 and §3.1] Some sentences are grammatically incomplete, e.g., ‘These inconsistencies also hinders the potential implementations using digital technologies (e.g.LLMs).’ The manuscript should be carefully proofread for such errors.
Circularity Check
No significant circularity: the prompt-condition effects are empirical, and the missing expert ground truth is a validity limitation rather than a circular reduction.
full rationale
The paper's claimed derivation chain is an experimental comparison: four prompt conditions with increasing levels of metric formalization and description are fed to GPT-4, and the resulting score distributions are compared statistically. The observed increase in concentration and consistency for Model-C4 is an empirical outcome of the prompts, not a quantity that is defined in terms of the outcome or fitted to it. The 'expert knowledge' inputs (metrics, quantifiers, descriptions) are collected from the literature and presented to the model before inference; they are not derived from the model's outputs, and no parameter is fitted to the evaluation data. The paper does overstate the conclusion by calling higher consistency 'enhanced evaluative performance' without a human-expert or objective walkability benchmark; the Conclusion itself defers practitioner comparison to future work. That is a construct-validity or overclaim concern, not circularity: there is no equation or definitional reduction showing that the measured consistency is the same thing as the claimed performance. Self-citations (e.g., the Triple-A ontology [16] and metric-structuring work [4, 15]) are used only to organize and describe metrics; they are not load-bearing for the central empirical claim that altering prompt specificity changes score distributions. The observation that Model-C4's scoring descriptions include restrictive rules (e.g., presence-based binary scores) is acknowledged in the paper and explains part of the variance reduction, but this is reported as a prompt-design effect rather than as a predetermined or renamed input. No specific circular step satisfying the required quote-and-reduction standard can be exhibited, so the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (2)
- C1 overall score scale (1 to 105)
- Per-metric score scale (1 to 5) for C2-C4
assumptions (4)
- domain assumption The walkability metrics taken from the two review articles [9,12] are valid expert knowledge for evaluating safety and attractiveness.
- domain assumption The Triple-A ontology categorisation provides an appropriate formal structure for the metrics.
- domain assumption ChatGPT-4 can reliably perceive street-view images and correctly apply the provided metrics.
- domain assumption The selected 42 SVIs from KartaView are sufficient and representative enough to compare prompt conditions.
Cite this review
Pith. "Pith review of Can a Large Language Model Assess Urban Design Quality? Evaluating Walkability Metrics Across Expertise Levels." pith.science (2026). https://pith.science/paper/JJ3NCIVT
@misc{pith2026250421040,
author = {Pith},
title = {Pith review of: Can a Large Language Model Assess Urban Design Quality? Evaluating Walkability Metrics Across Expertise Levels},
year = {2026},
howpublished = {\url{https://pith.science/paper/JJ3NCIVT}},
note = {Machine review of arXiv:2504.21040}
}
read the original abstract
Urban street environments are vital to supporting human activity in public spaces. The emergence of big data, such as street view images (SVIs) combined with multimodal large language models (MLLMs), is transforming how researchers and practitioners investigate, measure, and evaluate semantic and visual elements of urban environments. Considering the low threshold for creating automated evaluative workflows using MLLMs, it is crucial to explore both the risks and opportunities associated with these probabilistic models. In particular, the extent to which the integration of expert knowledge can influence the performance of MLLMs in evaluating the quality of urban design has not been fully explored. This study sets out an initial exploration of how integrating more formal and structured representations of expert urban design knowledge into the input prompts of an MLLM (ChatGPT-4) can enhance the model's capability and reliability in evaluating the walkability of built environments using SVIs. We collect walkability metrics from the existing literature and categorize them using relevant ontologies. We then select a subset of these metrics, focusing on the subthemes of pedestrian safety and attractiveness, and develop prompts for the MLLM accordingly. We analyze the MLLM's ability to evaluate SVI walkability subthemes through prompts with varying levels of clarity and specificity regarding evaluation criteria. Our experiments demonstrate that MLLMs are capable of providing assessments and interpretations based on general knowledge and can support the automation of multimodal image-text evaluations. However, they generally provide more optimistic scores and can make mistakes when interpreting the provided metrics, resulting in incorrect evaluations. By integrating expert knowledge, the MLLM's evaluative performance exhibits higher consistency and concentration.
Figures
Figures from the paper (4 more)
Reference graph
Works this paper leans on
-
[1]
arXiv preprint arXiv:2303.08774 (2023)
Achiam, J., Adler, S., Agarwal, S., Ahmad, L., Akkaya, I., Aleman, F.L., Almeida, D., Altenschmidt, J., Altman, S., Anadkat, S., et al.: Gpt-4 technical report. arXiv preprint arXiv:2303.08774 (2023)
arXiv 2023
-
[2]
Transport reviews 40(2), 183–203 (2020)
Arellana, J., Saltar´ ın, M., Larra˜ naga, A.M., Alvarez, V., Henao, C.A.: Urban walkability considering ped- estrians’ perceptions of the built environment: a 10-year review and a case study in a medium-sized city in latin america. Transport reviews 40(2), 183–203 (2020)
work page 2020
-
[3]
Ariffin, R.N.R., Rahman, N.H.A., Zahari, R.K.: Systematic literature review of walkability and the build environment. J. Pol’y & Governance 1, 1 (2021)
work page 2021
-
[4]
Ataman, C., Herthogs, P., Tun¸ cer, B., Perrault, S.: Multi-criteria decision making in digital participation (2022)
work page 2022
-
[5]
Bender, E.M., Gebru, T., McMillan-Major, A., Shmitchell, S.: On the dangers of stochastic parrots: Can language models be too big? In: Proceedings of the 2021 ACM conference on fairness, accountability, and transparency. pp. 610–623 (2021)
2021
-
[6]
Landscape and Urban Planning 215, 104217 (2021)
Biljecki, F., Ito, K.: Street view imagery in urban analytics and gis: A review. Landscape and Urban Planning 215, 104217 (2021)
work page 2021
-
[7]
Trunfio, G.: Enhancing urban walkability assessment with multimodal large language models
Bleˇ ci´ c, I., Saiu, V., A. Trunfio, G.: Enhancing urban walkability assessment with multimodal large language models. In: International Conference on Computational Science and Its Applications. pp. 394–411. Springer (2024)
work page 2024
-
[8]
doi.org/10.6084/m9.figshare.28236869.v1 (2025)
Cai, C.: MLLM assessments for walkability. doi.org/10.6084/m9.figshare.28236869.v1 (2025)
Show all 32 references
-
[9]
Applied Sciences 13(7), 4408 (2023)
Dragovi´ c, D., Krkljeˇ s, M., Slavkovi´ c, B., Aleksi´ c, J., Radakovi´ c, A., Ze´ cirovi´ c, L., Alcan, M., Hasanbegovi´ c, E.: A literature review of parameter-based models for walkability evaluation. Applied Sciences 13(7), 4408 (2023)
2023
-
[10]
Journal of Urban design 14(1), 65–84 (2009)
Ewing, R., Handy, S.: Measuring the unmeasurable: Urban design qualities related to walkability. Journal of Urban design 14(1), 65–84 (2009)
2009
-
[11]
Landscape ecology 33, 323–340 (2018)
Fan, P., Wan, G., Xu, L., Park, H., Xie, Y., Liu, Y., Yue, W., Chen, J.: Walkability in urban landscapes: A comparative study of four large cities in china. Landscape ecology 33, 323–340 (2018)
2018
-
[12]
International Journal of Sustainable Transportation 16(7), 660–679 (2022)
Fonseca, F., Ribeiro, P.J., Conticelli, E., Jabbari, M., Papageorgiou, G., Tondelli, S., Ramos, R.A.: Built en- vironment attributes and their influence on walkability. International Journal of Sustainable Transportation 16(7), 660–679 (2022)
2022
-
[13]
British journal of sports medicine 44(13), 924–933 (2010)
Frank, L.D., Sallis, J.F., Saelens, B.E., Leary, L., Cain, K., Conway, T.L., Hess, P.M.: The development of a walkability index: application to the neighborhood quality of life study. British journal of sports medicine 44(13), 924–933 (2010)
2010
-
[14]
The lancet 388(10062), 2912–2924 (2016)
Giles-Corti, B., Vernez-Moudon, A., Reis, R., Turrell, G., Dannenberg, A.L., Badland, H., Foster, S., Lowe, M., Sallis, J.F., Stevenson, M., et al.: City planning and population health: a global challenge. The lancet 388(10062), 2912–2924 (2016)
2016
-
[15]
Computers, Environment and Urban Systems 113, 102178 (2024)
Grisiute, A., Wiedemann, N., Herthogs, P., Raubal, M.: An ontology-based approach for harmonizing metrics in bike network evaluations. Computers, Environment and Urban Systems 113, 102178 (2024)
2024
-
[16]
Herthogs, P.: Triple-a design: a mid-level ontology for design goals and design evaluation (2021), 2021
2021
-
[17]
Buildings 15(1), 113 (2024)
Huang, G., Yu, Y., Lyu, M., Sun, D., Dewancker, B., Gao, W.: Impact of physical features on visual walkability perception in urban commercial streets by using street-view images and deep learning. Buildings 15(1), 113 (2024)
2024
-
[18]
Vintage Canada (2010)
Jacobs, J.: Dark age ahead: Author of the death and life of great American cities. Vintage Canada (2010)
2010
-
[19]
Journal of transport geography 112, 103646 (2023)
de Jong, T., Fyhri, A.: Spatial characteristics of unpleasant cycling experiences. Journal of transport geography 112, 103646 (2023)
2023
-
[20]
Journal of the American Heart Association 9(12), e016152 (2020)
Koohsari, M.J., Nakaya, T., Hanibuchi, T., Shibata, A., Ishii, K., Sugiyama, T., Owen, N., Oka, K.: Local- area walkability and socioeconomic disparities of cardiovascular disease mortality in japan. Journal of the American Heart Association 9(12), e016152 (2020)
2020
-
[21]
Transportation 46, 2347–2379 (2019)
Larranaga, A.M., Arellana, J., Rizzi, L.I., Strambi, O., Cybis, H.B.B.: Using best–worst scaling to identify barriers to walkability: A study of porto alegre, brazil. Transportation 46, 2347–2379 (2019)
2019
-
[22]
Journal of transport & health 18, 100880 (2020)
Lee, S., Lee, C., Nam, J.W., Abbey-Lambertz, M., Mendoza, J.A.: School walkability index: Application of environmental audit tool and gis. Journal of transport & health 18, 100880 (2020)
2020
-
[23]
Li, Z., Wang, Y., Song, Z., Huang, Y., Bao, R., Zheng, G., Li, Z.J.: What can llm tell us about cities? arXiv preprint arXiv:2411.16791 (2024)
2024 arXiv
-
[24]
Cities 150, 105022 (2024) Title Suppressed Due to Excessive Length 13
Liu, L., Sevtsuk, A.: Clarity or confusion: A review of computer vision street attributes in urban studies and planning. Cities 150, 105022 (2024) Title Suppressed Due to Excessive Length 13
2024
-
[25]
In: Proceedings of the 2nd ACM SIGSPATIAL International Workshop on Spatial Big Data and AI for Industrial Applications
Liu, X., Haworth, J., Wang, M.: A new approach to assessing perceived walkability: Combining street view imagery with multimodal contrastive learning model. In: Proceedings of the 2nd ACM SIGSPATIAL International Workshop on Spatial Big Data and AI for Industrial Applications....
2023
-
[26]
Computers, Environment and Urban Systems 117, 102243 (2025)
Malekzadeh, M., Willberg, E., Torkko, J., Toivonen, T.: Urban attractiveness according to chatgpt: Con- trasting ai and human insights. Computers, Environment and Urban Systems 117, 102243 (2025)
2025
-
[27]
International journal of environmental research and public health 11(1), 527–536 (2014)
Pelclov´ a, J., Fr¨ omel, K., Cuberek, R.: Gender-specific associations between perceived neighbourhood walkability and meeting walking recommendations when walking for transport and recreation for czech inhabitants over 50 years of age. International journal of environmental ...
2014
-
[28]
Pereira, M.F., Almendra, R., Vale, D.S., Santana, P.: The relationship between built environment and health in the lisbon metropolitan area–can walkability explain diabetes’ hospital admissions? Journal of Transport & Health 18, 100893 (2020)
2020
-
[29]
Bulletin of Geography
Reisi, M., Nadoushan, M.A., Aye, L.: Local walkability index: Assessing built environment influence on walking. Bulletin of Geography. Socio-economic Series (46), 7–21 (2019)
2019
-
[30]
Travel Behaviour and Society 39, 100983 (2025)
Wedyan, M., Saeidi-Rizi, F.: Assessing the impact of walkability indicators on health outcomes using machine learning algorithms: A case study of michigan. Travel Behaviour and Society 39, 100983 (2025)
2025
-
[31]
Computers, Environment and Urban Systems 64, 288–296 (2017)
Yin, L.: Street level urban design qualities for walkability: Combining 2d and 3d gis measures. Computers, Environment and Urban Systems 64, 288–296 (2017)
2017
-
[32]
Annals of the American Association of Geographers 114(5), 876–897 (2024)
Zhang, F., Salazar-Miranda, A., Duarte, F., Vale, L., Hack, G., Chen, M., Liu, Y., Batty, M., Ratti, C.: Urban visual intelligence: Studying cities with artificial intelligence and street-level imagery. Annals of the American Association of Geographers 114(5), 876–897 (2024)
2024
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.