{"id":"8eb41bd4-b57f-449b-b61f-3701eb850903","arxiv_id":"2504.21040","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"Giving ChatGPT-4 metric definitions improves consistency and concentration of walkability scores on street view images, but accuracy against expert judgment is not validated.","lead":"This paper tested whether giving ChatGPT-4 clearer expert rules changes how it scores street walkability from photos. On 42 Singapore street images, adding detailed metric definitions made scores more consistent, but no human expert benchmark was used.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central claim that expert knowledge enhances evaluative performance is unsupported because no human expert or objective ground truth is used; the paper's own Conclusion defers practitioner comparison to future work.","rationale":"The reader's weakest assumption identifies exactly the load-bearing gap: the study defines improved evaluative performance as higher consistency and concentration, with no human expert or objective benchmark to define correctness. The paper's own Conclusion lists the missing practitioner comparison as future work, which is a clear self-identified limitation. My analysis confirms this concern and adds two supporting confounds: the C1-vs-C2-C4 scale/aggregation mismatch, and the C2-vs-C3 metric-name confound. These do not overturn the descriptive findings—adding metrics and descriptions clearly shifts and concentrates scores—but they do undermine the stronger causal claim about enhanced performance. The reader's CONDITIONAL verdict is appropriate: the paper should be accepted only if the central claim is reframed as a descriptive effect, or if the planned human-expert validation is added. Since the reader already reached this conclusion, no verdict change is needed.","tokens_in":10173,"tokens_out":3196,"duration_ms":35408,"concrete_test":"Obtain expert walkability ratings for the same 42 SVIs from at least two urban design practitioners using the same 21 safety and 21 attractiveness metrics, then compute per-model agreement with the expert consensus (e.g., intraclass correlation, Kendall's tau, or mean absolute error) for Models-C1 through C4. If Model-C4's agreement is not significantly better than Model-C1/C2/C3, the central claim that expert knowledge enhances evaluative performance fails, even though the descriptive variance-reduction effect may remain.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's strongest claim—'MLLMs' evaluative performance can be enhanced by integrating expert knowledge'—rests on interpreting higher score concentration and consistency as improved performance. However, no ground truth is established: there is no human expert assessment, no objective walkability audit, and no external benchmark against which the models' scores are validated. The Conclusion explicitly lists 'engaging urban design practitioners to evaluate the SVIs and compare their assessment with the MLLMs' as future work, acknowledging that the missing comparison is essential. Without such a reference, the observed variance reduction in Model-C4 could reflect more rigid or biased scoring, not greater correctness. The paper does document concrete cases where Model-C4 avoids misinterpretations (e.g., not treating crosswalks as fixed furniture, Figure 6), but these are anecdotal and do not quantify overall accuracy. Additionally, the comparison between Model-C1 and Models-C2/C3/C4 is confounded: C1 uses a holistic 1–105 scale while C2–C4 sum 21 metrics each scored 1–5, so differences may partly arise from aggregation format rather than expert knowledge. Similarly, the C2-to-C3 comparison changes both metric wording (vague vs quantified names) and semantic clarity, making the specific effect of 'semantic clarity' ambiguous. The descriptive claim that prompt structure affects score distributions is supported by the statistical tests, but the evaluative-performance claim is not.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper asks whether a multimodal large language model (ChatGPT-4) can evaluate urban design quality from street-view images (SVIs), and whether injecting increasingly formal expert knowledge into the prompt changes the resulting scores. The authors collect 124 walkability metrics from the literature, select 21 safety and 21 attractiveness metrics, and construct four prompt conditions: C1 (no metrics, holistic 1–105 score), C2 (vague metric names, 1–5 per metric), C3 (quantified metric names, 1–5 per metric), and C4 (quantified metrics plus formal descriptions and scoring rules). They apply the four models to 42 SVIs from Singapore and analyze the score distributions with Levene’s test, Welch’s ANOVA, Games-Howell post-hoc tests, and Kruskal-Wallis tests. The main descriptive findings are that C1 produces more optimistic and dispersed scores, while C2–C4 are mutually closer in overall distribution, and that on selected metrics C4 yields more concentrated scores. Two example images illustrate cases where C4 avoids misinterpretations (e.g., not treating crosswalks as fixed furniture). The paper concludes that integrating expert knowledge enhances MLLMs’ evaluative performance and that increasing semantic clarity improves consistency.","tokens_in":10417,"tokens_out":3912,"duration_ms":41559,"significance":"If the descriptive findings hold, the paper provides a useful empirical demonstration that prompt structure and metric definition materially change MLLM-based street-environment scoring, and that well-specified rubrics reduce variance. The publication of the metric lists, prompts, and assessment data on Figshare is a reproducible resource for the urban-analytics community, and the use of inferential statistics rather than only point estimates is a strength. However, the central claim as worded in the Conclusion—that evaluative performance is enhanced—goes beyond what the data can show, because no human expert assessment or objective walkability benchmark is used. The manuscript itself defers practitioner comparison to future work. The significance of the paper therefore rests on the well-supported descriptive claims about score distributions and prompt sensitivity, not on the unsupported claim of improved correctness.","major_comments":[{"comment":"The claim that integrating expert knowledge enhances MLLMs’ evaluative performance is not supported by the evidence presented. The study measures consistency, concentration, and differences in score distributions; it does not measure correctness against any ground truth. The Conclusion explicitly lists ‘engaging urban design practitioners to evaluate the SVIs and compare their assessment with the MLLMs’ as future work, which acknowledges that the missing expert baseline is essential. Without such a reference, a more concentrated score distribution could reflect more rigid or systematically biased scoring, not better performance. The paper should either reframe the central claim as an effect on score distribution and consistency, or add a human-expert evaluation of the same SVIs to support the performance language.","section":"Conclusion and §1"},{"comment":"The comparison between Model-C1 and Models-C2/C3/C4 is confounded by the response scale: C1 uses a single holistic 1–105 score, while C2–C4 sum 21 metrics each scored 1–5 (also a 21–105 range, but with a different aggregation structure). Observed differences in means, variances, and rank order may therefore partly reflect the difference between holistic and decomposed scoring formats rather than the absence or presence of expert knowledge. To support the claim that expert knowledge changes evaluations, the authors should compare C1 against, for example, a decomposed but metric-free condition, or analyze normalized/standardized scores that make the two formats more comparable.","section":"§3.2, Table 3, and Figure 4"},{"comment":"The inference that Model-C4’s higher concentration reflects fewer misinterpretations is based on two anecdotal examples. The paper states that ‘the more varied and dispersed score distributions in Model-C3 and Model-C2 could be attributed to the ambiguity resulting from the lack of definitions,’ but this causal interpretation is not quantitatively tested. To make this load-bearing point credible, the authors should systematically classify the model’s reasoning across all 42 images per metric (for example, by coding each response as consistent or inconsistent with the provided description), rather than relying on two hand-picked cases.","section":"§4.1, Figures 5 and 6"},{"comment":"The statistical support for the central variance-related claim is partial. Table 3 shows that C2, C3, and C4 are not significantly different in overall safety and attractiveness distributions, yet the paper later emphasizes C4’s higher concentration. This is not necessarily contradictory because the metric-level analysis in Figure 5 is more fine-grained, but the paper should state clearly that the concentration effect is metric-specific and not a global property of Model-C4’s scores. The current presentation risks overgeneralizing the finding.","section":"§4.1 and Table 3"}],"minor_comments":[{"comment":"The text says the prompts draw on metrics ‘outlined in Section 5,’ but the metrics are presented in Section 3.1 and Table 1; the cross-reference should be corrected.","section":"§3.2"},{"comment":"The entries ‘DiverseLandscape LandscapeDiversityIndex’ and ‘Colorfulness EnvironmentalColorDiversity’ appear in both the vague and quantified columns for Safety and Attractiveness, which is confusing and may be a formatting artifact. The table should clearly distinguish the two sets of metric names or explain that some names coincide.","section":"Table 1"},{"comment":"The paper states that models were tested ‘in order from Model-C1 to Model-C4’ to prevent learning from descriptions, but it is not clear whether each test used a fresh session or whether the same conversation history was retained. If a single session was used, earlier prompts could still influence later responses. Please clarify the session and conversation-reset protocol.","section":"§3.2"},{"comment":"Some sentences are grammatically incomplete, e.g., ‘These inconsistencies also hinders the potential implementations using digital technologies (e.g.LLMs).’ The manuscript should be carefully proofread for such errors.","section":"§2 and §3.1"}],"recommendation":"major_revision","confidential_remarks":"The paper’s core descriptive finding—that prompt structure and metric formalization change MLLM score distributions—is publishable, but the title, abstract, and conclusion overclaim by using ‘evaluative performance’ without a human-expert or ground-truth baseline. The missing comparison is acknowledged in the paper itself as future work. A major revision that either adds such a comparison or substantially softens the performance claim would bring the manuscript in line with its evidence. The small sample (42 SVIs, one city, one model) is a further concern but is already acknowledged as a limitation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is an exploratory study of whether prompt-level 'expert knowledge' changes ChatGPT-4's walkability scoring of street view images. The genuinely new piece is the systematic four-level prompt comparison built on an ontology-structured metric list, with the metric database released. The descriptive result holds up: scores become more concentrated and consistent as metric definitions are added (Model C4), and the statistical tests are appropriate. Two concrete examples of misinterpretations avoided by the defined prompts are useful, though anecdotal.\n\nWhere the paper oversells is in the conclusion: 'MLLMs' evaluative performance can be enhanced by integrating expert knowledge.' That claim requires a ground truth, and there is none. No human expert assessment, no objective walkability audit, no benchmark. Consistency and concentration are not correctness. The authors themselves list practitioner comparison as future work, which is the right next step and should have been part of this study. There is also a scale confound: C1 uses a holistic 1-105 scale while C2-C4 sum 21 metrics of 1-5, so the C1 vs others difference may be partly aggregation format. The C2-to-C3 comparison changes both metric phrasing and semantic clarity, so the specific effect of clarity remains ambiguous. The sample of 42 images is small and single-run LLM outputs add noise.\n\nThese are meaningful soft spots, but they are not fatal to the paper's descriptive contribution. The authors are honest about limitations and the statistical work is sound. As an initial exploration, it deserves a serious referee. I would suggest asking the authors to either add a small human expert comparison or reframe the central claim as improved consistency and reduced misinterpretations, not enhanced evaluative performance.\n\nWho gets value: researchers working on MLLM-based urban analytics or street view evaluation, and planners considering automated walkability screening. Not a definitive study, but a useful stepping stone.\n\nMy recommendation: send to peer review, with requests for a toned-down claim and a human baseline.","headline":"A useful exploratory study of prompt-level expertise effects on MLLM walkability scoring, but the central evaluative-performance claim needs a human ground truth it currently lacks.","tokens_in":10961,"tokens_out":1941,"would_cite":false,"duration_ms":18571,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Giving a multimodal language model formal, expert-written walkability metric definitions in its prompts makes its ratings of street-view images more consistent and less optimistic.","keywords":["walkability assessment","street view images","multimodal large language models","expert knowledge","prompt design","urban design quality","pedestrian safety","attractiveness"],"falsifier":"Have expert urban designers score the same street-view images with the same 21 safety and 21 attractiveness metrics, then compare their scores with the model's four prompt conditions after normalizing the different score scales; if human expert scores do not agree more with the fully described condition than with the no-metrics condition, the claim that expert-knowledge prompts improve evaluative performance fails. A second check is to run the same four prompts on a different multimodal model: if the concentration effect disappears, it is a property of the model rather than of expert-knowledge prompting.","tokens_in":9962,"feed_emoji":"🚶","tokens_out":8964,"duration_ms":85271,"temperature":0.7,"pith_summary":"This paper asks whether a multimodal large language model can be steered toward expert-grade evaluation of street walkability by changing how expert knowledge is written into the prompt. Using ChatGPT-4 on 42 street-view images from Singapore, the authors compare four prompt conditions: no metrics, vague metric names from the literature, metric names with quantifiers, and quantified metric names plus formal descriptions and scoring rules. They find that the no-metrics condition gives systematically optimistic scores and can rank the same street differently from the metric-informed conditions, while the fully described condition produces the most concentrated, consistent per-metric distributions and avoids specific misinterpretations such as counting a crosswalk as fixed furniture. The pith is that the form and semantic clarity of expert knowledge in a prompt, not just its presence, change and stabilize what the model reports, making prompt design a lever for automated urban design evaluation.","feed_headline":"Expert rules in prompts stabilize AI street walkability scores","feed_subtitle":"Four prompt levels scored 42 Singapore street images; formal metric definitions cut optimistic and mistaken AI ratings.","key_machinery":"The carrying mechanism is a four-level expertise gradient built into the prompt: Level 1 asks for safety and attractiveness ratings with no metrics, Level 2 uses vague metric names from the literature, Level 3 uses quantified metric names, and Level 4 adds a formal description and scoring rule for each metric. The metric set itself is assembled from two walkability review literatures and structured through an ontology-based categorisation, with 21 safety metrics and 21 attractiveness metrics selected as comparison sets. The formal descriptions are the active ingredient: they convert short labels into operational scoring instructions, reducing the ambiguity that lets the model read crosswalks as fixed furniture or infer traffic calming devices from narrow roads. Statistical tests (Levene, Welch ANOVA, Games-Howell, Kruskal-Wallis) are used to show that the prompt conditions produce different distributions and that the fully described condition concentrates scores.","core_discovery":"The paper's central claim is that a multimodal large language model's evaluative performance can be enhanced by integrating expert knowledge, and that increasing the semantic clarity of that knowledge improves consistency of the evaluative outputs. Concretely, when ChatGPT-4 is given no evaluation criteria it rates pedestrian safety and attractiveness more optimistically and diverges from metric-informed models; when given literature metrics, overall score distributions shift and stabilize; and when given quantified metric names plus formal descriptions with scoring rules, per-metric scores become more concentrated and the model stops making certain interpretive errors. The authors do not claim these expert-informed scores are objectively correct: comparison with human urban design practitioners is explicitly left to future work.","pith_inferences":["If the consistency gain is later confirmed against human expert raters, the same prompt-document technique could transfer to other perceptual urban qualities such as enclosure, imageability, or maintenance without retraining the model.","Because the overall score distributions of the three metric-informed conditions showed no significant differences, the practical payoff of formal descriptions may lie in variance reduction and error correction rather than in shifting average scores; that distinction deserves an explicit test.","A natural extension is to measure inter-run reliability by repeating each prompt condition several times, since the current single-run design cannot separate prompt-induced concentration from the model's sampling noise.","With one model and one city's images, the safest reading is that expert-knowledge prompting changes this model's behaviour on this dataset; how far the effect extends across models, languages, and street networks remains open."],"forward_implications":["Automated walkability screening can be steered without fine-tuning: writing formal metric descriptions into the prompt materially changes how the model rates street-view images.","Unaided multimodal language model ratings are optimism-prone, so any automated urban quality workflow that skips metric definitions should treat its scores cautiously.","Formal descriptions reduce per-metric misinterpretation, making scores more concentrated and better comparable across streets within a metric.","Quantified metric names alone do not significantly shift overall score distributions; the descriptive layer is what aligns the model with the intended criteria at the metric level.","The approach can translate low scores on design-actionable metrics into targeted urban design interventions, demonstrated for two lower-scoring streets in the study."],"supporting_citations":[{"why":"Supplies one of the two review literatures from which the 124 walkability metrics and the safety/attractiveness criteria were collected.","marker":"[9]"},{"why":"Supplies the second review source for built-environment walkability attributes and their measurements.","marker":"[12]"},{"why":"Provides the ontology-based structuring used to categorise and organise the collected metrics into comparable classes.","marker":"[16]"},{"why":"Establishes the expert urban design qualities that motivate converting criteria into measurable walkability metrics.","marker":"[10]"},{"why":"Documents the GPT-4 model that all four prompt conditions were tested on.","marker":"[1]"},{"why":"Shows prior use of multimodal large language models for walkability assessment, the immediate line of work this study extends.","marker":"[7]"}],"fun_headline_variants":["Expert prompts make AI walkability scores more consistent","AI street walkability ratings stabilize with expert rule prompts","ChatGPT-4 urban design scores get steadier with expert metrics","Adding expert rules to AI prompts tightens walkability scoring","Formal metrics in prompts curb AI's over-optimistic street ratings"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that higher consistency and concentration of the model's scores count as better evaluative performance, since the paper includes no human expert scoring or objective walkability benchmark to confirm that the more concentrated, expert-informed scores are actually right.","fun_headline_variants_meta":{"raw":{"variants":["Expert prompts make AI walkability scores more consistent","AI street walkability ratings stabilize with expert rule prompts","ChatGPT-4 urban design scores get steadier with expert metrics","Adding expert rules to AI prompts tightens walkability scoring","Formal metrics in prompts curb AI's over-optimistic street ratings"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000195,"raw_usage":{"total_tokens":1374,"prompt_tokens":976,"completion_tokens":398,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":592,"completion_tokens_details":{"reasoning_tokens":315}},"tokens_in":592,"tokens_out":398,"duration_ms":4243,"temperature":1.0,"reasoning_tokens":315,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T05:46:24.198747+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Have expert urban designers score the same street-view images with the same 21 safety and 21 attractiveness metrics, then compare their scores with the model's four prompt conditions after normalizing the different score scales; if human expert scores do not agree more with the fully described condition than with the no-metrics condition, the claim that expert-knowledge prompts improve evaluative performance fails. A second check is to run the same four prompts on a different multimodal model: if the concentration effect disappears, it is a property of the model rather than of expert-knowledge prompting.","supporting_citations":[{"cited_title":"Applied Sciences 13(7), 4408 (2023)","cited_arxiv_id":null,"evidence_quote":"Supplies one of the two review literatures from which the 124 walkability metrics and the safety/attractiveness criteria were collected."},{"cited_title":"International Journal of Sustainable Transportation 16(7), 660–679 (2022)","cited_arxiv_id":null,"evidence_quote":"Supplies the second review source for built-environment walkability attributes and their measurements."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the ontology-based structuring used to categorise and organise the collected metrics into comparable classes."},{"cited_title":"Journal of Urban design 14(1), 65–84 (2009)","cited_arxiv_id":null,"evidence_quote":"Establishes the expert urban design qualities that motivate converting criteria into measurable walkability metrics."},{"cited_title":"Trunfio, G.: Enhancing urban walkability assessment with multimodal large language models","cited_arxiv_id":null,"evidence_quote":"Shows prior use of multimodal large language models for walkability assessment, the immediate line of work this study extends."}],"review_version":1}