{"id":"fcb3e7fa-8f7e-4062-8637-ef85b054df6b","arxiv_id":"2505.08034","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A user-centric survey of generative AI applications for smart cities, covering conversational interfaces for citizens, operators, and planners.","lead":"This paper surveys how generative AI chatbots and tools are being used in smart cities, organized around three user groups: citizens, operators, and planners. It is a useful orientation for anyone tracking how cities are starting to adopt large language models.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Novelty claim of being the first user-centric GenAI smart-city survey is unsupported: no search methodology or differentiation from prior surveys, despite the paper's own hedging.","rationale":"The reader's weakest assumption concerns the unproven taxonomy and non-transparent example selection. My concern is closely related but focuses on the firstness claim itself: the survey's contribution is explicitly framed as the first user-centric synthesis, and this depends on a systematic absence of prior work, which the paper neither demonstrates nor even describes how to verify. The paper's own hedged wording strengthens the concern, as does the inclusion of non-conversational systems under a \"conversational interfaces\" scope. This does not require rejecting the paper; a survey can still be useful as a thematic synthesis. However, the conditional acceptance already given is the right level: the novelty claim needs either supporting evidence or softening, and the review methodology should be stated. I therefore recommend no change to the reader's verdict, and I agree only partially with the reader's identified weakest assumption because the reader emphasized taxonomy representativeness and example selection rather than the explicit firstness claim.","tokens_in":12531,"tokens_out":3907,"duration_ms":39200,"concrete_test":"Run a systematic literature search in IEEE Xplore, ACM DL, Scopus, and arXiv up to May 2025 with a documented query combining \"generative AI\" or \"large language model\" with \"smart city\" and at least one of \"citizen\", \"operator\", \"planner\", \"user-centric\", or \"conversational\". Screen the retrieved surveys to determine whether any prior work organizes GenAI/LLM applications by user role or explicitly covers the same three archetypes. In particular, inspect the paper's own cited surveys [9], [16], [23], [31], [43], and [58]. If a prior survey is found, the \"first\" claim is false and should be softened or qualified; if none is found, record the search protocol as evidence supporting the claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central contribution is the claim, in the Introduction and abstract, that this is \"the first paper ... that talks about GenAI applications in the context of the three main user archetypes\" and \"the first comprehensive summarization of GenAI techniques for Smart Cities from the lens of the critical users.\" The load-bearing condition is that no earlier survey already organizes GenAI/LLM smart-city applications by user role or by conversational interfaces. The paper supplies no search protocol, inclusion criteria, database coverage, or explicit differentiation from prior surveys such as Xu et al. [9], Zhang et al. [16], Salierno et al. [23], Wang et al. [31], Feng et al. [25], and Xu et al. [43]. The prose itself hedges: \"to the best of our knowledge\" appears in Section I and \"We believe\" in the abstract, acknowledging that firstness has not been verified. Because the survey's value as a \"first synthesis\" rests on this novelty, the claim is currently unsupported rather than demonstrated. A related internal tension also bears on the claim: the stated scope is \"conversational interfaces,\" yet several operator/planner entries are non-conversational predictive or synthetic-data systems (e.g., LLMAir [45], STLLM [59], UrbanGPT [8], PlacemakingAI [51]), so the distinctive \"user-centric conversational\" framing is partially diluted. The fix is not to abandon the claim but to either substantiate it with a systematic, reproducible review or explicitly reframe the paper as a non-exhaustive thematic survey.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a survey of generative AI (GenAI) applications in smart cities, organized around three user archetypes: citizens, operators/managers, and urban planners. It reviews the foundational technologies (IoT, Digital Twins, GenAI), presents example applications for each archetype, and discusses challenges and future directions. The authors claim that this is the first paper to survey GenAI applications in smart cities from the perspective of these three user archetypes, with a particular focus on conversational interfaces built on urban data foundations.","tokens_in":12785,"tokens_out":2279,"duration_ms":22670,"significance":"If the firstness claim and the survey's coverage are substantiated, the paper would provide a useful synthesis for researchers and practitioners: the user-archetype lens is a sensible organizing principle, the catalog of recent systems (UrbanGPT, CityGPT, IncidentResponseGPT, VayuBuddy, ACQAR, etc.) is timely, and the discussion of RAG and human-in-the-loop mitigation for hallucination reflects current practice. The paper also names concrete pitfalls (e.g., the NYC MyCity chatbot failure). However, the significance is conditional on verifying the novelty claim through a reproducible search protocol and on correcting the citation of performance statistics that currently come from a secondary source. The survey does not contain derivations or machine-checked proofs, but that is not a weakness for a survey; its value rests on coverage, accuracy, and framing.","major_comments":[{"comment":"The central claim that this is \"the first paper to the best of our knowledge\" and \"the first comprehensive summarization\" is unsupported: the manuscript provides no literature search protocol, no database coverage, no inclusion/exclusion criteria, and no explicit differentiation from prior surveys that overlap substantially, such as Xu et al. [9], Zhang et al. [16], Salierno et al. [23], Wang et al. [31], Feng et al. [25], and Xu et al. [43]. Without a systematic, reproducible methodology or a concrete comparison showing how the present survey is distinct from these existing works, the firstness claim remains an assertion rather than a demonstrated contribution.","section":"Section I and Abstract"},{"comment":"The performance statistics for city deployments are not backed by primary sources: the 94% user satisfaction rate for Barcelona, the 42% increase in first-time resolution and 28% cost reduction for Vienna are all attributed to Ref. [23], which is a secondary encyclopedia article, and no independent verification is cited. Because these numbers are used as evidence that conversational GenAI yields measurable operational benefits, the paper should either cite the original city reports or explicitly flag these as secondary-source claims that require verification.","section":"Section III.A"},{"comment":"The stated scope of the survey is \"conversational interfaces,\" but several entries in the operator and planner sections are non-conversational predictive or synthetic-data systems: LLMAir [45] performs air-quality prediction, STLLM [59] is an edge-computing PM2.5 forecaster, UrbanGPT [8] is a spatiotemporal traffic-flow predictor, and PlacemakingAI [51] and the Land Use Configuration GAN [31] are generative visualization tools rather than conversational interfaces. The paper should either justify how these systems fit under the conversational-interface framing (e.g., as backend components of a conversational assistant) or narrow the stated scope so the survey's coverage matches its framing.","section":"Section III.B and III.C"},{"comment":"The selection of Citizens, Operators, and Planners as \"the three critical user archetypes\" is asserted without supporting evidence or a discussion of how the archetype taxonomy was derived, and the paper does not explain how representative examples were chosen for each group. A survey's usefulness depends on the representativeness of its examples; the authors should provide a brief rationale for the taxonomy and for the application-selection process, or acknowledge the selection as illustrative rather than systematic.","section":"Section III introduction"}],"minor_comments":[{"comment":"Reference [52] contains a typo: \"WWang\" should be \"Wang.\"","section":"References"},{"comment":"The sentence on Vienna's results is ambiguous: \"a 42% increase [23] in first-time resolution\" should specify whether this is a 42% increase in the first-time resolution rate or a 42% increase in the number of cases resolved on first contact.","section":"Section III.A"},{"comment":"The text reads \"UrbanGPT UrbanGPT [8]\" with a duplicated model name; one occurrence should be removed.","section":"Section III.B"},{"comment":"In the sentence \"Large Language Models (LLMs) [14] - based on Generative Pre-trained Transformers (GPT) [9]\", the citation [9] appears to be a generic GenAI/urban digital twin reference rather than the foundational GPT reference; this citation should be corrected or removed.","section":"Section II.D"},{"comment":"The manuscript contains several minor typographical spacing issues (e.g., \"S mart\", \"s trategic\" in the abstract and body) that should be cleaned up before publication.","section":"Various"}],"recommendation":"major_revision","confidential_remarks":"The paper appears to be an IEEE conference manuscript. The firstness claim in the abstract and introduction is a strong marketing statement that the authors hedge with \"to the best of our knowledge\" and \"We believe\"; without a systematic review protocol, this claim could be easily challenged by a reviewer or reader familiar with the cited surveys. The citation of performance figures from a secondary encyclopedia source is a factual-accuracy risk that should be fixed before publication. The self-citations [21] and [56] are peripheral and do not appear to distort the survey's conclusions, but they should be kept in proportion. If the authors can reframe the contribution as a user-centric organizing synthesis rather than a strict firstness claim, and correct the sourcing issues, the paper would be acceptable for a conference venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nQuick take on arXiv:2505.08034. It's a competent, clearly organized survey that uses a genuinely helpful device — grouping GenAI-for-smart-city applications by three user archetypes: citizens, operators, planners. That frame is the paper's real contribution, and it works: it makes the sprawling literature easier to navigate and surfaces a sensible set of recent systems (UrbanGPT, CityGPT, eGridGPT, VayuBuddy, ACQAR) organized around what each user actually needs. The writing is plain, the figures are fine, and the challenges section (hallucination, cost, bias, HITL) is balanced.\n\nThe soft spots are in the novelty and methodology, not the content. The authors claim this is 'the first paper ... that talks about GenAI applications in the context of the three main user archetypes' and 'the first comprehensive summarization of GenAI techniques for Smart Cities from the lens of the critical users.' They hedge with 'to the best of our knowledge' and 'We believe,' which is honest, but the claim still outruns what they show. There's no search protocol, no inclusion criteria, no database coverage, and no explicit comparison against earlier surveys (several of which — Xu et al. [9], Zhang et al. [16], Salierno et al. [23] — already cover large parts of this ground). A quick sanity check suggests the three-archetype framing may indeed be less common, but without a systematic sweep it's an assertion, not a demonstrated result.\n\nSecond, some of the headline statistics (Barcelona 94% satisfaction, Vienna 42% first-time resolution) come from a single secondary encyclopedia source. That doesn't make them wrong, but a survey should flag them as secondhand or chase the primary report.\n\nThird, a scope tension: the abstract and intro promise 'conversational interfaces,' yet a number of included systems (LLMAir, STLLM, UrbanGPT) are non-conversational predictive or synthetic-data models. That dilutes the distinctive framing and makes the selection feel a bit arbitrary. The two self-citations are peripheral and not a problem.\n\nThe issues are fixable. Either do a proper, reproducible literature search and defend the firstness claim against prior surveys, or drop the 'first' language and present this as a thematic review. I'd send it out for review — it has enough substance and a useful organizing idea — but with the clear expectation of a revision before acceptance.\n\nFor your own work: the taxonomy is worth borrowing even if the survey itself is not the last word.","headline":"Useful user-centric taxonomy in a competent survey, but the 'first' claim and missing methodology need fixing.","tokens_in":13337,"tokens_out":2517,"would_cite":true,"duration_ms":23328,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that conversational generative AI can make smart-city data usable by three distinct audiences—citizens, operators, and planners—and surveys the systems that already point in that direction.","keywords":["generative AI","smart cities","large language models","conversational interfaces","urban digital twins","Internet of Things","synthetic data generation","user-centric taxonomy"],"falsifier":"A systematic census of GenAI smart-city applications—starting from the paper's own citations and a broad literature search—that counts the intended user for each would settle the claim; finding a substantial share serving roles outside the three archetypes, such as emergency first responders or private developers, would falsify the organizing taxonomy.","tokens_in":12336,"feed_emoji":"🏙️","tokens_out":10349,"duration_ms":92413,"temperature":0.7,"pith_summary":"Smart cities generate more data from IoT sensors, official records, and digital twins than most people can actually use. This survey argues that conversational generative AI can close that gap by letting users ask questions in plain language and receive grounded, situation-aware answers. Its organizing claim is that GenAI applications in cities serve three user archetypes: citizens seeking services and everyday information, operators managing live urban systems, and planners testing long-term scenarios. The paper reviews deployed and proposed systems for each group and connects them to the existing data foundation of city records, IoT streams, and Urban Digital Twins. It positions itself as the first comprehensive survey of GenAI for smart cities organized from this user-centric perspective.","feed_headline":"Three user roles organize GenAI smart-city apps","feed_subtitle":"Conversational AI, sorted by who uses it: citizens, operators, and planners.","key_machinery":"The load-bearing structure is the three archetypes—Citizens, Operators/Managers, and Planners—each with distinct information needs and interaction styles. The mechanism that carries the argument is the conversational interface built on Large Language Models, grounded by Retrieval-Augmented Generation in city-controlled documents or IoT and Digital Twin data, and guarded by human-in-the-loop review where factual accuracy is critical. The paper uses this machinery to sort a wide range of proposed and deployed systems into a coherent map, and to show how each archetype can draw on the common city data foundation of official records, sensor streams, and Urban Digital Twins.","core_discovery":"The paper's central discovery is that the same underlying data foundation can serve three distinct urban roles through natural-language interfaces, and that the triad of Citizens, Operators, and Planners is the right lens for organizing GenAI applications in smart cities. For citizens, it finds deployed city-service chatbots with reported gains such as higher satisfaction, faster first-contact resolution, and lower service costs, plus proposals for transit, routing, and air-quality assistance. For operators, it identifies LLM-based systems that summarize incidents, explain anomalies, and support grid, water, and traffic management by grounding answers in real-time data. For planners, it collects simulation and synthetic-data tools that generate traffic scenarios, visualize proposed spaces, and test what-if policies before real-world implementation. The paper claims to be the first comprehensive summarization of these techniques from this user-centric viewpoint.","pith_inferences":["The three-archetype taxonomy may understate other consequential users of urban data, such as emergency first responders, private developers, tourists, and civic technologists; a census of actual deployments would show whether the frame is too narrow.","The city-reported metrics cited in the paper—for example 94% satisfaction, 42% first-contact resolution, and 28% cost reduction—could be assembled into a shared benchmark for future citizen-facing GenAI deployments, though the paper itself does not standardize them.","If retrieval grounding and human-in-the-loop review remain the core safeguards, the same deployment pattern likely applies beyond cities to any public-sector organization putting LLMs in front of official documents."],"forward_implications":["Cities that adopt the three-archetype framing can organize GenAI strategy into three tracks: citizen-facing grounded chatbots, operator-facing real-time interrogation and anomaly explanation, and planner-facing simulation and synthetic scenario tools.","The cited deployments indicate that retrieval-augmented generation and human review are the main practical safeguards against hallucination; without them, factual errors can break trust, as the New York City chatbot episode cited in the paper shows.","Existing municipal data—official records, IoT streams, and Urban Digital Twins—is treated as sufficient raw material for first-generation applications, so near-term pilots do not require new city infrastructure.","Synthetic data generation is a corollary opportunity: planners can test policies and operators can fill sensor gaps in simulation before committing to real-world changes."],"supporting_citations":[{"why":"supplies the scoping-review base for using GenAI to generate urban data, scenarios, and 3D city models on Urban Digital Twins.","marker":"[9]"},{"why":"describes a Smart City Digital Twin framework used to illustrate multi-source data integration challenges.","marker":"[13]"},{"why":"provides the deployed city examples and satisfaction or resolution metrics for citizen-facing LLM services.","marker":"[23]"},{"why":"contributes CityGPT, a planner-facing tool for querying urban knowledge graphs and spatiotemporal mobility patterns.","marker":"[25]"},{"why":"contributes the UrbanLLM foundational-model proposal for citizen trip planning and urban activity management.","marker":"[26]"},{"why":"shows a retrieval-augmented public-transit assistant grounding schedules and real-time data for citizen queries.","marker":"[33]"},{"why":"documents the human-in-the-loop deployment that mitigates hallucination in citizen question answering.","marker":"[36]"},{"why":"is the Urban Generative Intelligence reference, cited for the foundational platform integrating LLMs with a city simulator and knowledge graph.","marker":"[43]"},{"why":"grounds operator-facing applications by retrieving data for energy-infrastructure digital twins.","marker":"[49]"}],"fun_headline_variants":["Conversational GenAI for citizens, operators, planners","Smart cities: AI that talks, sorted by user role","First user-centric survey of urban GenAI apps","GenAI meets city roles: citizen, operator, planner"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The survey's organization rests on the claim that Citizens, Operators, and Planners are the three user types that matter for GenAI in smart cities, and the paper offers no systematic evidence that these categories cover the applications that actually exist.","fun_headline_variants_meta":{"raw":{"variants":["Conversational GenAI for citizens, operators, planners","Smart cities: AI that talks, sorted by user role","First user-centric survey of urban GenAI apps","GenAI meets city roles: citizen, operator, planner"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000211,"raw_usage":{"total_tokens":1391,"prompt_tokens":901,"completion_tokens":490,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":517,"completion_tokens_details":{"reasoning_tokens":426}},"tokens_in":517,"tokens_out":490,"duration_ms":5019,"temperature":1.0,"reasoning_tokens":426,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T22:04:28.822724+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A systematic census of GenAI smart-city applications—starting from the paper's own citations and a broad literature search—that counts the intended user for each would settle the claim; finding a substantial share serving roles outside the three archetypes, such as emergency first responders or private developers, would falsify the organizing taxonomy.","supporting_citations":[{"cited_title":"UrbanLLM: Autonomous Urban Activity Planning and Management with Large Language Models","cited_arxiv_id":null,"evidence_quote":"contributes the UrbanLLM foundational-model proposal for citizen trip planning and urban activity management."},{"cited_title":"Generative AI Assistant for Pub lic Transport Using Scheduled and Real-Time Data","cited_arxiv_id":null,"evidence_quote":"shows a retrieval-augmented public-transit assistant grounding schedules and real-time data for citizen queries."},{"cited_title":"A Retrieval-Augmented Genera tion Approach for Data-Driven Energy Infrastructure Digital Twins","cited_arxiv_id":null,"evidence_quote":"grounds operator-facing applications by retrieving data for energy-infrastructure digital twins."}],"review_version":1}