{"id":"cc64dafc-87e4-45ab-8431-5e9a88255c87","arxiv_id":"2606.11678","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.5,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"On UPBench, 25 LLMs show a non-monotonic planning curve—strong Remember/Analyze, weak Understand/Evaluate—with four failure modes that support differential, not blanket, AI delegation.","lead":"Researchers built UPBench, a planning-specific test of 25 LLMs across four knowledge pillars and five cognitive levels. Models often look strong on analysis yet fail on precise conceptual understanding and jurisdiction-specific judgment, which matters for how agencies should use AI.","discovery_kind":"new_application","skeptic_critique":{"model":"grok-4.5","headline":"Non-monotonic curve and structural-limit claims rest on unvalidated item-to-cognition mapping and moderate dual-track agreement without human baselines.","rationale":"The reader correctly isolates the dual-track protocol and exam-derived items as the weakest assumption supporting the strongest claim. My concern is the same load-bearing point, sharpened by the paper’s own internal signals (abstract/body mismatch on the curve, Figure 4 corrections, ρ=0.67, n=5 panel, no human baseline). These do not invalidate the directional usefulness of UPBench or the differential-delegation takeaway, but they keep absolute percentages and the “structural rather than developmental” reading of phronetic deficit under-supported. Hence the verdict remains CONDITIONAL; no upgrade or downgrade is warranted until the proposed human-baseline check (or equivalent construct-validation) is run. Agreement with the reader is therefore full on the critical assumption.","tokens_in":26144,"tokens_out":543,"duration_ms":7054,"concrete_test":"Have an independent panel of ≥10 licensed planners (AICP or equivalent, balanced US/China experience) score a stratified 60-item subset (12 per cognitive level) under the same rubrics used for Track 2; compute human mean accuracy and inter-rater reliability per level, then re-plot the cognitive curve. If human Understand accuracy is also depressed relative to Remember/Analyze, or if human–LLM rank-order correlation on the four diagnostics is low, the structural-limit interpretation weakens.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim (non-monotonic cognitive curve + four epistemic diagnostics as structural AI limits on planning phronesis) requires that the 405 scenarios, after cross-national equivalence mapping (§3.2) and Bloom tagging, validly elicit the intended cognitive levels and that dual-track scores (LLM-as-judge + 5-expert panel, final Spearman ρ=0.67, Table 1) measure professional judgment rather than exam-style pattern matching. The paper itself shows abstract–body inconsistency on the curve (abstract: better on higher-order than factual recall; body §4.2: Remember 89.6% high, Understand 55.3% collapse, Analyze 81.8% recovery), large original-vs-verified pillar corrections (Figure 4, up to −17.4% on Practice), and no reported human-planner baseline on the same items. Without those anchors, the U-shape and diagnostics (regulatory hallucination etc.) may overstate structural incapacity relative to developmental or construct-validity artifacts.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper introduces UPBench, a 4×5 domain-specific benchmark (four planning knowledge pillars × five Bloom-adapted cognitive levels) for evaluating whether LLMs can reason like professional urban planners. Using 405 bilingual (US/China) scenarios, a dual-track protocol (LLM-as-judge for lower levels; five-expert panel for Analyze/Evaluate), and evaluation of 25 models, the authors report a non-monotonic performance curve across cognitive levels, pillar asymmetries, cross-national origin effects, and four epistemic failure modes (regulatory hallucination, conceptual conflation, wickedness paralysis, phronetic deficit). They argue for differential delegation of planning tasks to AI rather than wholesale substitution, with implications for practice and education.","tokens_in":26459,"tokens_out":1608,"duration_ms":20077,"significance":"The work fills a clear gap: planning lacks a theory-grounded, profession-specific AI evaluation framework comparable to Med-PaLM or LegalBench. Strengths include scale (25 models, 405 scenarios), explicit grounding in phronesis and plan-quality evaluation traditions, dual-track scoring with reported calibration (Table 1), qualitative appendices with verbatim failure examples, and a practical differential-delegation framing. Code is released. If the non-monotonic curve and diagnostics hold under stronger construct validation, the paper would be a durable reference for AI governance in planning agencies and for reorienting planning education toward institutional literacy and normative judgment.","major_comments":[{"comment":"Abstract vs. §4.2 inconsistency on the central empirical claim. The abstract states models “perform better on higher-order analytical tasks than on factual recall and integrative judgment.” §4.2 and Figure 2 report the opposite ordering for factual recall: Remember 89.6% (highest), Understand 55.3% (collapse), Apply 76.2%, Analyze 81.8%, Evaluate 37.9%. The body’s U-shape (high Remember, low Understand, recovery at Analyze, low Evaluate) is the load-bearing result; the abstract misstates it. Align abstract, takeaway, and discussion with the reported means before any claim about “inverted” or “non-monotonic” gradients is used for differential delegation.","section":null},{"comment":"Construct validity of cognitive-level operationalization, especially Evaluate (§3.1). Remember/Understand/Apply map reasonably to item formats, but Evaluate is defined as “precise retrieval of domain-specific terminology” (e.g., supplying “Multiple nuclei model”). That is closer to Remember/Understand than to Bloom’s Evaluate (criterion-based judgment) or to phronesis. If Evaluate items are terminology fill-ins, the 37.9% floor and the “phronetic deficit” diagnosis partly reflect item design, not structural incapacity for professional judgment. Re-map or re-label levels, or replace Evaluate items with genuine multi-criteria judgment tasks, and re-estimate the curve.","section":null},{"comment":"No human-planner baseline on the same 405 items. Claims that the non-monotonic curve and four diagnostics mark structural AI limits on planning phronesis (§4.2–4.4, §5.1) require an anchor: how do licensed planners or advanced students score under the same dual-track rubrics? Without that, low Understand/Evaluate scores may reflect hard or poorly calibrated items rather than AI-specific failure. Report at least a small human baseline (or pilot) on a stratified subset and discuss relative gaps.","section":null},{"comment":"Figure 4 and data integrity. Figure 4 shows original pillar means underestimated by up to −17.4% (Planning Practice) relative to “verified” data, compressing the pillar range from a claimed 48.2%–72.4% to 65.6%–69.3%. The manuscript does not explain how the original numbers were produced, what was corrected, or whether Table 2 / cognitive-level means were similarly revised. Clarify the verification pipeline and ensure all reported aggregates (including CN/US splits and the non-monotonic curve) come from a single audited scoring run.","section":null},{"comment":"Dual-track validity for “professional judgment” (§3.3, Table 1). Final automated–expert Spearman ρ = 0.67 after nine prompt iterations is only moderate; the expert panel is n = 5; Track 1 uses LLM-as-judge for Remember/Understand/Apply. The paper treats this as acceptable by analogy to plan-quality intercoder ranges, but the central claim is about professional judgment, not plan completeness. Report inter-expert agreement (not only judge–expert), sensitivity of the U-curve to Track 1 vs Track 2, and whether conclusions change if only expert-scored Analyze/Evaluate cells are used for the phronesis claims.","section":null}],"minor_comments":[{"comment":"§3.1: Bloom operationalization for Evaluate is internally inconsistent with the later claim that Evaluate “most directly tests phronesis.” Fix the prose so level definitions match the theoretical claims.","section":null},{"comment":"Table 2 is split across pages and repeats the header awkwardly; consider a single compact table or supplementary full matrix with CN/US averages.","section":null},{"comment":"GitHub is named PlanBench while the paper uses UPBench; align naming to avoid confusion with related PlanGPT work.","section":null},{"comment":"Several typos and spacing artifacts (e.g., “an non-monotonic,” “remain remain,” “domain-specificfine-tuned,” missing spaces after periods in §1–2). Full copy-edit pass needed.","section":null},{"comment":"§2.3 claims UPBench addresses “all five deficits” but only three are listed immediately above; fix the count.","section":null},{"comment":"Figure 1 caption says 405 items each for CN and EN; confirm whether total is 405 or 810 and make N consistent throughout.","section":null},{"comment":"Cross-national equivalence mapping (§3.2) is described at a high level; a short appendix table of example Chinese→US instrument mappings would strengthen reproducibility.","section":null}],"recommendation":"major_revision","confidential_remarks":"Important and timely paper for a planning/urban-informatics or interdisciplinary AI venue; the phronesis framing is a genuine contribution. The abstract–body mismatch and Figure 4 verification gap are fixable but currently undermine trust in the headline result. I would not accept until the cognitive-level construct (especially Evaluate), human baseline (even small), and audited aggregates are cleaned up. Scope fits journals that publish plan-quality methods and professional-education work as well as AI-benchmark papers; less ideal for pure ML venues that demand human baselines and inter-annotator stats as default."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The thing to know: this is a serious, usable domain benchmark for planning, not another generic MMLU clone. UPBench’s 4×5 matrix (four pillars from licensure/curricula × five Bloom levels), bilingual US–China equivalence mapping, 25 models, 405 scenarios, dual-track scoring, and four named failure modes (regulatory hallucination, conceptual conflation, wickedness paralysis, phronetic deficit) are real contributions. The differential-delegation takeaway is practical and well-matched to the data.\n\nWhat works: they actually ran a large model set, published calibration (Table 1, best Spearman ρ≈0.67 after nine prompt iterations), qualitative appendices with item-level cases, and a clear origin effect (US models favor US scenarios; China-origin models are more balanced). Cross-disciplinary synthesis and Analyze-level narrative assembly look relatively strong; Evaluate and jurisdiction-specific Apply look weak. That pattern is worth citing for AI-in-planning and professional-education debates. Code pointer is there.\n\nSoft spots, in proportion. The abstract claims models do better on higher-order tasks than factual recall; §4.2 and the numbers show Remember ~89.6% high, Understand collapsing to ~55.3%, then Apply/Analyze recovering, Evaluate low. That is a non-monotonic U, not “higher-order beats recall.” Figure 4 openly shows large original-vs-verified pillar corrections (up to −17.4% on Practice), so early narrative overstated domain gaps. Expert panel is n=5; no human-planner baseline on the same items; ρ=0.67 is moderate and they treat it as acceptable by plan-quality standards, which is fair but not strong. The stress-test is right that item-to-cognition mapping and exam-style sources can inflate “structural phronetic limits” relative to developmental or construct artifacts—but the paper still shows coherent failure modes and institutional sensitivity, so the directional claim holds even if the absolute percentages and “inherently human” rhetoric should be dialed back.\n\nMath/data/citations: empirical, not formal; tables and appendices are inspectable; literature (Flyvbjerg, Schön, LegalBench/Med-PaLM, plan quality) is engaged honestly; related PlanGPT work is cited without circular scoring.\n\nWho it is for: planning scholars, AI-for-professions people, and agencies writing AI use policies. Worth a serious referee. I would engage, cite the matrix and diagnostics with the curve caveat, and not treat the structural-limit language as settled.","headline":"Useful first theory-grounded planning LLM bench with a real multi-model pattern, but the abstract misstates the curve and the structural-limit claim outruns the construct validation.","tokens_in":27086,"tokens_out":622,"would_cite":true,"duration_ms":8691,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"LLMs excel at planning analysis but fail at institutional facts and practical judgment, so agencies should use them only under differential delegation.","keywords":["artificial intelligence","planning knowledge","professional judgment","phronesis","benchmark","large language models","planning education","differential delegation"],"falsifier":"If frontier models or domain-fine-tuned systems reverse the non-monotonic curve—matching or exceeding human experts on Understand and Evaluate items that require jurisdiction-specific regulatory application and normative commitment—while expert panels still rate the reasoning as professionally adequate, the claim of structural (not merely developmental) limits collapses.","tokens_in":27031,"feed_emoji":"🏙️","tokens_out":894,"duration_ms":10923,"temperature":0.7,"pith_summary":"This paper asks which parts of professional urban planning knowledge large language models can actually replicate. The authors build UPBench, a 4x5 test matrix covering four knowledge pillars (principles, cross-disciplinary integration, governance, practice) and five cognitive levels from Remember to Evaluate, then score 25 models on 405 bilingual scenarios drawn from Chinese and U.S. planning materials. The central result is a non-monotonic performance curve: models score high on Remember and Analyze yet collapse on Understand and Evaluate. The authors read this pattern as evidence that planning’s supposedly “basic” knowledge is densely institutional and jurisdictional, so pattern-matching fails where human planners rely on situated, value-laden judgment. They codify the failures into four diagnostics and translate them into a practical rule of differential delegation—use AI for synthesis and first drafts, keep humans for regulation, norms, and context-sensitive procedure.","feed_headline":"AI plans well analytically, fails on zoning facts and judgment","feed_subtitle":"25 models show a non-monotonic curve; use them only under differential delegation","key_machinery":"UPBench: a bilingual 4×5 matrix of four knowledge pillars (Principles of Urban Planning, Cross-Disciplinary Integration, Planning Governance, Planning Practice) by five Bloom-adapted cognitive levels, scored by a dual-track protocol of LLM-as-judge plus expert panel.","core_discovery":"Across 25 LLMs and 405 UPBench scenarios, planning competence is non-monotonic by cognitive level (Remember ~89.6%, Understand ~55.3%, Apply ~76.2%, Analyze ~81.8%, Evaluate ~37.9%). Models handle broad analytical synthesis better than precise conceptual understanding or integrative judgment, because planning’s “lower-order” knowledge is institutionally and temporally embedded. The authors formalize the resulting limits as four epistemic diagnostics—regulatory hallucination, conceptual conflation, wickedness paralysis, and phronetic deficit—and conclude that AI should be delegated only where these modes do not dominate.","pith_inferences":["The same institutional-density argument likely applies to other phronetic professions (architecture, public administration, social work) whose “basic” knowledge is jurisdictionally encoded.","If the Understand collapse is architectural rather than data-limited, continued scaling alone will not close the gap; retrieval-augmented or institution-specific systems may still leave the normative and phronetic failures intact.","Longitudinal re-runs of UPBench on successive model generations would cleanly test whether the four diagnostics are temporary or structural ceilings."],"forward_implications":["Agencies can productively use LLMs for literature review, scenario generation, and preliminary cross-disciplinary analysis under ordinary professional review.","Any AI-assisted regulatory interpretation or procedural advice requires structured human verification against current local law.","Planning education should shift emphasis from transmitting facts AI already approximates toward institutional literacy, normative courage, and critical evaluation of AI outputs.","Training-data origin strongly shapes cross-national performance, so “global” planning models will systematically under-serve underrepresented institutional systems."],"fun_headline_variants":["LLMs excel at planning analysis but fail zoning facts and judgment","25 models: non-monotonic skill, strong synthesis weak on institutional recall","AI handles analysis better than factual or evaluative planning tasks","Models ace higher analysis yet struggle with embedded planning knowledge","Planning competence non-monotonic: analysis high, judgment and facts low"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That the dual-track scoring of adapted licensure and curriculum items, after calibration to moderate expert agreement, genuinely measures professional planning judgment rather than exam-style pattern matching.","fun_headline_variants_meta":{"raw":{"variants":["LLMs excel at planning analysis but fail zoning facts and judgment","25 models: non-monotonic skill, strong synthesis weak on institutional recall","AI handles analysis better than factual or evaluative planning tasks","Models ace higher analysis yet struggle with embedded planning knowledge","Planning competence non-monotonic: analysis high, judgment and facts low"]},"model":"grok-4.5","effort":"low","cost_usd":0.004282,"raw_usage":{"total_tokens":1359,"prompt_tokens":872,"num_sources_used":0,"completion_tokens":87,"cost_in_usd_ticks":42820000,"prompt_tokens_details":{"text_tokens":872,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":400,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":872,"tokens_out":87,"duration_ms":5225,"temperature":1.0,"reasoning_tokens":400,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-14T18:04:11.082042+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"If frontier models or domain-fine-tuned systems reverse the non-monotonic curve—matching or exceeding human experts on Understand and Evaluate items that require jurisdiction-specific regulatory application and normative commitment—while expert panels still rate the reasoning as professionally adequate, the claim of structural (not merely developmental) limits collapses.","supporting_citations":[],"review_version":2}