{"id":"a974a1c0-affb-4424-8464-d9c7a803145d","arxiv_id":"2501.08019","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Using XGBoost, SHAP, and counterfactual explanations, the study generates minimal building-height and setback changes that improve Sky View Factor and visibility in synthetic urban blocks with 4 to 6 percent average simulation error.","lead":"This paper presents a machine learning framework that predicts how city parks score on sky openness and visibility, then suggests small changes to surrounding buildings to improve them. The authors report that their counterfactual approach finds workable adjustments about 15 to 30 times faster than a genetic algorithm benchmark.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"CFX vs GA comparison is not apples-to-apples: offline training cost omitted and final solution quality never compared.","rationale":"The reader's weakest_assumption focuses on generalizability from synthetic Tehran-style blocks to diverse urban contexts. While that is a real limitation, the single most load-bearing premise of the central claim is the quantitative efficiency comparison: CFX beats GA by 15–30x. That comparison is currently unfair because it omits the offline cost of building the surrogate and never reports the actual optimization quality achieved by GA. A reviewer cannot verify the headline 'significantly faster' claim without these numbers. This is not an internal inconsistency in the mathematics, but a missing control in the experimental design. The paper still has independent support: the held-out Grasshopper validation of CFX recommendations (RMSE 4–6%) is a genuine positive, and the SHAP-based feature analysis is coherent. The concern is fixable by adding the missing measurements, which is why the reader's CONDITIONAL verdict remains appropriate. I would not escalate to REJECT because the framework may still be practically valuable even if the speed advantage is smaller when amortized; but the current presentation overstates the evidence. My agreement is partial because the reader's rationale mentions the unfair GA comparison, but the formal weakest_assumption field pointed to generalizability rather than this missing control.","tokens_in":15794,"tokens_out":6958,"duration_ms":71459,"concrete_test":"On the same 10 configurations used in Figure 5, run both GA and CFX from identical initial blocks with the same objective (e.g., maximize SVF or raise visibility class). Record: (i) final SVF/visibility value achieved by each method, (ii) wall-clock time for GA to reach that value, (iii) wall-clock time for the full CFX pipeline including the 1,152 simulations and XGBoost training. If GA's best objective exceeds CFX's by more than 5% relative, or if CFX total pipeline time exceeds 15 minutes, the abstract's 'significantly faster performance' and 'comparable' claims require revision.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central efficiency claim (Abstract; §4.4) compares CFX's ~1 minute to GA's 15–30 minutes, but this is not a like-for-like comparison. The GA time is the wall-clock cost of simulation-based optimization (50 individuals × 3–4 generations, each 6–9 s per evaluation). The CFX time is only the counterfactual-search step on an already-trained surrogate; it excludes the 1,152 Grasshopper runs used to generate the training set (§3.1) and the XGBoost training/hyperparameter selection (§3.2). Moreover, §4.4 never reports the actual SVF/visibility values achieved by the GA—it only says 'closest results' were reached—so 'comparable convergence' is asserted, not demonstrated. If the offline data-generation and training time is included, or if GA actually reaches a substantially better SVF/visibility than CFX on the same scenarios, the headline speed advantage and 'similar outcomes' claim collapse. This is the load-bearing premise of the paper's contribution.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper presents a workflow for suggesting minimal, localized changes to urban open-space layouts to improve sky view factor (SVF) and visibility. The authors generate 1,152 synthetic urban block configurations based on simplified Tehran morphology, simulate SVF and visibility with Ladybug/Grasshopper, train five surrogate machine-learning models, select XGBoost as the most accurate, analyze feature importance with SHAP, and generate counterfactual explanations (CFX) via a KD-tree search. The CFX suggestions are re-simulated in Grasshopper and achieve mean RMSE between 4.12% and 6.06%. The paper claims that the CFX approach is 15–30 times faster than a genetic algorithm benchmark while producing comparable results.","tokens_in":15991,"tokens_out":5115,"duration_ms":48939,"significance":"If the speed and accuracy claims hold, the framework would be a useful interpretable surrogate-based design-support tool for early-stage urban retrofitting. The study includes a genuine held-out model comparison (XGBoost SVF R²=0.91; visibility F1 up to 0.89) and a validation of counterfactual suggestions against simulation, which is a good internal consistency check. The integration of SHAP and CFX on a simulation-based surrogate is a meaningful contribution in this applied context. However, the paper's central efficiency claim is not yet supported by the evidence, as discussed in the major comments. No code or dataset is provided, so reproducibility could not be verified.","major_comments":[{"comment":"The headline speed comparison is not like-for-like. The reported GA time of 15–30 minutes is wall-clock simulation time for a population of 50 over 3–4 generations, whereas the reported CFX time of about 1 minute covers only the counterfactual retrieval step on an already-trained surrogate; the 1,152 Grasshopper simulations used to generate the training set (§3.1) and the XGBoost training and hyperparameter selection (§3.2) are excluded. In addition, the text states that GA reached 'closest results' but never reports the actual SVF or visibility values achieved by GA on the same scenarios, so the claim of 'similar outcomes' is not demonstrated. To support the speed advantage, the comparison should include the full pipeline cost for both methods and report the achieved objective values of the GA solutions.","section":"§4.4"},{"comment":"The validation of CFX suggestions by Grasshopper simulations is not an independent check on accuracy because the same Ladybug/Grasshopper pipeline generated the training labels in §3.1; it only confirms that the surrogate does not grossly diverge from the simulation engine. The abstract and conclusion claim suitability for 'real-world' urban planning, but no comparison against measured SVF/visibility data is provided. Please add a discussion of this limitation or include field or independent simulation data.","section":"§3.4 and Figure 5"},{"comment":"The claims of generalizability across diverse urban contexts are not supported by the experimental design. The 1,152 configurations are generated from a narrow set of rules (regular grid, no vegetation, fixed FAR values of 4.5 and 6.5, three orientations, two street widths, and 3–10 story heights), all based on a simplified Tehran morphology. Feature-importance rankings and CFX recommendations may therefore be specific to this distribution; testing on an independent urban dataset or a broader generative distribution is needed before claiming generalizability.","section":"§3.1 and §5"}],"minor_comments":[{"comment":"The notation in Tables 4 and 5 is unclear: empty cells, '0', and values such as 'dN 23 -14' make it difficult to see which features are changed and by how much per strategy. Please add explicit deltas or arrows (e.g., '−14 m') and fill all cells.","section":"Tables 4 and 5"},{"comment":"The Shapley formula in §3.3 is incorrectly typeset; the factorial term and the set difference are garbled. Please rewrite it in standard form: φ_i = Σ_{S⊆N∖{i}} (|S|!(n−|S|−1)!/n!) [v(S∪{i})−v(S)].","section":"Equation (1)"},{"comment":"The text for visibility SHAP values contains a direction typo: it says 'southeast (hSW)' when hSW denotes southwest. Please check directional abbreviations throughout the SHAP section.","section":"§4.2"},{"comment":"The choice of five counterfactuals per scenario is not justified; a sensitivity analysis showing how the number of counterfactuals affects solution quality would strengthen the optimization claims.","section":"§4.3"},{"comment":"The box plot in Figure 5 is described but not shown in the manuscript text; please include the figure or state that it is available in supplementary materials.","section":"§4.4 and Figure 5"},{"comment":"Some references contain typographical errors ('Chen et al., 2012s', 'Gomez-rodriguez' capitalization); please ensure the reference list is cleaned.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"This manuscript fits the journal's scope. The main risk is the fairness of the GA vs CFX timing comparison; I would ask the authors to re-benchmark with total pipeline cost and to report solution quality. If they cannot, the central claim is unsupported. The paper would also benefit from a statement about data/code availability."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The genuinely new thing here is the packaging: SHAP for feature importance plus counterfactual explanations (CFX) for minimal-change design suggestions, applied to SVF and visibility in urban open spaces, with the CFX outputs checked against Grasshopper simulations. That validation step is real work, and the mean RMSE of 4.12-6.06% across ten configurations is credible evidence that the counterfactuals are not just model hallucinations. The surrogate models are also evaluated on a held-out split, with XGBoost getting R2=0.91 for SVF and F1 up to 0.89 for visibility. No circularity in the target-result sense: the ML model is fit to simulation data, and the CFX suggestions are validated by independent simulations. The authors deserve credit for that. The soft spot is the one the stress-test flagged, and it is load-bearing. The paper's central claim is that CFX gets the same results as GA in about one minute versus 15-30 minutes. But the GA time is wall-clock simulation-based optimization, while the CFX time excludes the 1,152 Grasshopper runs used to build the training set and the XGBoost training and tuning. That is not apples-to-apples. Worse, Section 4.4 never reports the actual SVF or visibility values the GA achieved - it only says 'closest results' were reached in the 3rd or 4th generation. So 'comparable convergence' is asserted, not demonstrated. If the GA actually reaches substantially better values, the speed advantage is irrelevant. This is a fixable but serious flaw: the authors need to report final objective values for both methods, and either include full pipeline time for CFX or clearly scope the claim to inference-time only. Other concerns are minor by comparison. The synthetic scenarios are simplified Tehran-style blocks with no vegetation, fixed FAR ranges, and regular grids, so the generalizability claim in the title is overstated. Lack of code, data, and hyperparameter details would also need addressing. The SHAP interpretation sections read a bit long, but they are not wrong. Overall, this is a coherent applied study, not a breakthrough. The core idea - using counterfactuals for localized design retrofits - builds directly on Regenwetter et al. and existing ML-surrogate workflows. But the demonstration is honest and the validation is a plus. With a proper GA comparison and full timing accounting, the paper could be solid. As is, the speed claim is unsubstantiated, but the underlying methodology deserves a serious referee who can push for that revision.","headline":"A legitimate applied extension of SHAP+CFX to SVF/visibility optimization, with honest simulation validation, but the headline speed claim rests on an unfair GA comparison that needs fixing.","tokens_in":711,"tokens_out":950,"would_cite":false,"duration_ms":19945,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Park redesign answers arrive in 1 minute, not 30","keywords":["urban open spaces","Sky View Factor","visibility analysis","counterfactual explanations","XGBoost","SHAP","urban morphology","genetic algorithm optimization"],"falsifier":"Take real parks with surrounding building data, apply the CFX-recommended height and distance changes, and compare predicted SVF and visibility against fisheye photographs or high-fidelity re-simulation; if mean RMSE leaves the reported 4–6% band, or if a genetic algorithm beats CFX on solution quality when both are given the same ten-minute budget, the central claim would be refuted.","tokens_in":15612,"feed_emoji":"🏙️","tokens_out":7429,"duration_ms":70143,"temperature":0.7,"pith_summary":"The paper claims that urban open spaces can be optimized through small, targeted changes to surrounding buildings rather than full redesigns. It focuses on two morphology-driven metrics: sky view factor, the fraction of open sky seen from a park, and visibility, how well a park is seen from surrounding buildings. The authors train machine-learning models on 1,152 simulated Tehran-style blocks, find XGBoost most accurate, and then use counterfactual explanations to propose minimal edits such as lowering an eastern building two stories or shifting a southern neighbor farther away. Grasshopper re-simulation of the recommendations gives mean errors between 4.12% and 6.06%, and the full optimization runs in about a minute versus 15–30 minutes for a genetic algorithm. If true, planners gain a fast, interpretable way to retrofit existing open spaces without expensive global search.","feed_headline":"Park redesign answers arrive in 1 minute, not 30","feed_subtitle":"AI counterfactual search finds minimal building changes that open up parks, matching slow genetic searches at about 5% error.","key_machinery":"The engine is the counterfactual explanation, a small set of feature changes that flips a model's prediction to a target outcome. The implementation uses a KD-tree, a binary space-partitioning structure, to organize the parameter space of building heights and distances and quickly find the nearest feasible configuration that meets the goal, prioritizing features identified by SHAP and respecting fixed constraints such as orientation and street width. This mechanism replaces the iterative evaluate-and-mutate loop of a genetic algorithm with a single nearest-neighbor-style query on the trained XGBoost surrogate, which is what produces the speed advantage.","core_discovery":"The central discovery is that counterfactual explanation can serve as an optimizer for localized urban morphology. Using XGBoost as a surrogate for Ladybug/Grasshopper simulation, SHAP values to prioritize features, and a KD-tree search to locate the nearest feasible design change, the framework generates minimal adjustments that raise predicted SVF by up to 13 percentage points or lift a park's visibility class by one level. The paper reports that these recommendations are accurate: Grasshopper re-simulation of ten test configurations yields mean RMSE between 4.12% and 6.06%, while the CFX run takes about one minute compared with 15–30 minutes for a genetic algorithm. The claim is that explanation machinery, not global search, is sufficient for practically useful local improvements, with the extra benefit that each suggestion is directly actionable and interpretable.","pith_inferences":["Not directly claimed by the paper: the speed of CFX points to interactive design tools where dragging a building edge updates predicted sky view and visibility instantly, with KD-tree queries replacing repeated simulations.","Not directly claimed by the paper: the same surrogate-plus-counterfactual recipe should extend to other morphology-sensitive metrics such as daylight autonomy, wind comfort, or noise whenever enough simulated samples exist to train an accurate predictor.","Not directly claimed by the paper: because RMSE grows when CFX recommends large height or distance changes, adding uncertainty bands to counterfactual suggestions would shore up the method exactly where it is weakest.","Not directly claimed by the paper: a head-to-head test on real urban blocks with vegetation, irregular parcels, and setbacks would reveal whether the SHAP-dominant features remain dominant outside regular-grid synthetic configurations."],"forward_implications":["A designer can test five alternative localized retrofit strategies per park in about a minute before committing to detailed simulation.","Feature-importance rankings translate directly into design heuristics: park area and east/west building heights dominate SVF, while distance to southern buildings and building width dominate visibility.","Because SVF and visibility depend only on morphology rather than climate, the same trained pipeline can be reapplied to other cities without re-tuning to local weather, provided block geometry stays within the training envelope.","The 15–30x speed advantage over genetic algorithms makes block-by-block, city-scale optimization feasible on ordinary hardware.","The validated RMSE band of roughly 4–6% supports using CFX suggestions as a screening step, with high-fidelity simulation reserved for the most invasive recommended changes."],"supporting_citations":[{"why":"Supplies the KD-tree counterfactual search procedure that generates minimal design changes and drives the speed gain.","marker":"Alfeo et al., 2023"},{"why":"Establishes model-agnostic counterfactuals as a design-recommendation method, the conceptual basis for applying CFX to urban layouts.","marker":"Regenwetter et al., 2023"},{"why":"Frames counterfactual explanations in decision-making and strategic behavior, grounding the claim that CFXs give actionable change sets.","marker":"Tsirtsis and Gomez-rodriguez, 2020"},{"why":"Provides the genetic algorithm and NSGA-II baseline against which CFX speed and practicality are benchmarked.","marker":"Saad and Araji, 2021"},{"why":"Represents the prior genetic-algorithm urban-block optimization approach that the paper argues is too slow for localized adjustments.","marker":"Xu et al., 2019"},{"why":"Represents the parametric global-optimization framework that motivates the need for lower-cost localized alternatives.","marker":"Toutou et al., 2018"},{"why":"Defines sky view factor and its estimation methods, the target metric the framework optimizes.","marker":"Miao et al., 2020"},{"why":"Supplies the Shapley-value equation used to compute SHAP feature importances.","marker":"Nourkojouri et al., 2021"},{"why":"Connects SHAP values to counterfactual explanations, the combined interpretability recipe the framework adopts.","marker":"Zhong et al., 2022"},{"why":"Provides the SHAP game-theoretic attribution method used for feature importance assessment.","marker":"Slack et al., 2020"}],"fun_headline_variants":["AI park fixes: 1 minute vs 30","One-minute AI suggests minimal park changes","Counterfactual AI speeds park redesign 30x","Explainable AI finds small park tweaks fast","Local AI beats global search for urban parks"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that 1,152 randomly generated regular-grid blocks built from Tehran-style rules—no vegetation, fixed floor-area ratios, straight streets, and 3-to-10-story buildings—represent real urban open spaces well enough that recommendations trained on them remain valid in other contexts.","fun_headline_variants_meta":{"raw":{"variants":["AI park fixes: 1 minute vs 30","One-minute AI suggests minimal park changes","Counterfactual AI speeds park redesign 30x","Explainable AI finds small park tweaks fast","Local AI beats global search for urban parks"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000652,"raw_usage":{"total_tokens":2997,"prompt_tokens":957,"completion_tokens":2040,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":573,"completion_tokens_details":{"reasoning_tokens":1970}},"tokens_in":573,"tokens_out":2040,"duration_ms":15111,"temperature":1.0,"reasoning_tokens":1970,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T20:30:22.724884+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take real parks with surrounding building data, apply the CFX-recommended height and distance changes, and compare predicted SVF and visibility against fisheye photographs or high-fidelity re-simulation; if mean RMSE leaves the reported 4–6% band, or if a genetic algorithm beats CFX on solution quality when both are given the same ten-minute budget, the central claim would be refuted.","supporting_citations":[{"cited_title":"and Mohamed, W","cited_arxiv_id":null,"evidence_quote":"Represents the parametric global-optimization framework that motivates the need for lower-cost localized alternatives."}],"review_version":1}