{"id":"04fdaedd-e9f0-49e9-ad32-ac4cb9cfaa27","arxiv_id":"2501.11283","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"LLM agents can be wired to commercial radio-planning software to generate radio maps and run network optimization from a few prompts, but the reported coverage gains stem from the software's built-in optimizer.","lead":"The paper presents a software platform in which large language model agents automate radio map generation and wireless network planning by driving the commercial tool RANPLAN through natural language prompts. The demo shows fewer manual operations and improved coverage after automatic cell optimization, though the gains come largely from the optimizer rather than the language model.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Autonomy claim lacks reliability evidence: no error feedback, no success-rate evaluation, and only two curated demos.","rationale":"The reader's weakest_assumption identifies exactly this: the reliability of task planning, tool invocation, and log processing without human intervention. My analysis agrees and finds concrete textual support: the regex-based parsing in Section II-A3 and the explicit admission in Section IV that error feedback will only be added in the future. This is the most load-bearing concern because if the tool-call loop is not robust, the claimed 4-prompt automation fails in practice, and the two successful demonstrations are insufficient to establish a general capability. The confounded performance comparison (coverage/SINR improvements due to RANPLAN's built-in ACO) is a real but secondary issue: it undermines the claim that the LLM agent itself enhances coverage, but it does not directly attack the automation claim. My proposed test would provide the missing reliability evidence. Since the reader already set the verdict to CONDITIONAL, and this concern reinforces that conditionality, I leave the verdict unchanged.","tokens_in":6208,"tokens_out":4226,"duration_ms":42736,"concrete_test":"Run the two scenarios repeatedly (e.g., 50 independent trials per scenario) with identical prompts and record the end-to-end success rate, defined as completing all four tasks without human intervention. Additionally, inject perturbations: rephrase prompts, make the OSM server temporarily unavailable, or corrupt the downloaded file, and measure whether the agent recovers or fails. If the success rate is below, say, 95% or any perturbation leads to a hard failure, the autonomy claim is not supported. A complementary check: collect 100 LLM log outputs from varied user prompts and count how many are correctly parsed by the regular expression described in Section II-A3.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is that LLM agents autonomously generate radio maps and plan networks, reducing manual operations to four prompts. For this to hold, the agent must reliably map prompts to correct tool invocations and parse returned logs without human help. Section II-A3 describes using a regular expression to extract tool parameters from the LLM's log file; any formatting deviation would break the pipeline, and no fallback is provided. Section IV states that 'error feedback will be added' as future work, indicating the current system has no mechanism to detect or recover from tool-call failures. The experiments in Section III-B are two curated, successful runs only; no failure rate, no prompt variation, and no robustness tests are reported. The closing claim that this 'makes it possible to automate the radio map generation and wireless network planning in any locations on the earth' is therefore unsupported. Without quantitative evidence that the tool-call loop succeeds reliably, the automation claim reduces to a two-example demonstration rather than a general capability.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes an LLM-agent framework that drives the commercial RANPLAN Academic V6.8.0 software to generate radio maps and plan wireless networks from natural-language prompts. The framework comprises Profile, Tool, and Model modules plus memory and task-planning functionalities. Experiments on two real-world scenarios, HITSZ and Hyde Park, report that the agent reduces manual operations from 57 clicks to 4 prompts and improves the percentage of area with PL > -100 dB and SINR > 5 dB. The paper concludes that the approach makes it possible to automate radio map generation and wireless network planning in any location on Earth.","tokens_in":6300,"tokens_out":4783,"duration_ms":40519,"significance":"The work is a timely proof-of-concept for LLM-based control of a commercial network-planning tool. The operation-count comparison is concrete, easily checked, and accompanied by a working executable platform, which are clear strengths. If the baseline and reliability issues were resolved, the results would be useful to practitioners seeking to lower the entry barrier for network planning. As presented, however, the reported coverage/SINR gains are confounded with RANPLAN's built-in ACO optimizer, and the autonomy claim rests on only two curated runs with no error-rate or robustness analysis, so the demonstrated contribution is more modest than the abstract claims.","major_comments":[{"comment":"The baseline labeled 'without LLM' in Fig. 7 is the initial BS placement before the ACO step in RANPLAN, not a manual operator executing the same planning workflow without an LLM. Because the 'with LLM' runs include RANPLAN's built-in ACO, the reported improvements in PL and SINR percentages are attributable to the ACO algorithm, not to the LLM agent's task planning or memory. To support the abstract claim of enhanced coverage and SINR, the control condition should be the same workflow executed manually (including ACO), or an additional condition with the LLM agent but with ACO disabled. As reported, Fig. 7 does not cleanly separate the effect of the agent from the effect of the optimizer.","section":"Section III-B, Figs. 4, 6, and 7"},{"comment":"The agent's tool invocation depends on a single regular expression that parses the LLM's log file, as described in Section II-A3. If the LLM deviates from the expected formatting, the pipeline has no fallback, retry, or recovery mechanism, and the conclusion explicitly states that 'error feedback will be added' only in the future. Since the central claim is autonomous operation, the paper should provide evidence that the tool-call loop succeeds reliably across varied prompts, malformed outputs, or unexpected software states. The two curated scenarios in Section III-B are not sufficient for this, and the closing claim of automated operation 'in any locations on the earth' is unsupported without such reliability evidence.","section":"Section II-A3 and Section IV"},{"comment":"The derivation of the 91.4% and 93.0% reduction figures needs to be made explicit. Summing the 'without any LLM agents' row gives 57 operations, and four prompts yield a 93.0% overall reduction; the 91.4% figure appears to use only the first three columns (6+9+20) for radio map generation. The sentence 'reduces 91.4% and 93.0% manual operations for radio map generation and network planning tasks, respectively' is inaccurate because the network-planning column alone gives 95.5%, not 93.0%. Please clarify which rows and columns correspond to each task and define what counts as one manual operation.","section":"Section III-B and Table I"}],"minor_comments":[{"comment":"The phrasing 'often require complex manual operations' and 'due to heavy manual operations' repeats the same idea; please rephrase for clarity.","section":"Abstract"},{"comment":"There is a typo 'shot-term' for 'short-term', and a missing space in 'complete.The long-term memory'; please correct.","section":"Section II-B1"},{"comment":"Please use the standard capitalization 'PySide6' and 'Python' instead of 'PYSIDE 6' and 'PYTHON'.","section":"Section III-A"},{"comment":"The phrase 'statistical data' in connection with Fig. 7 is misleading because no error bars or multiple runs are reported; 'percentage data' would be more accurate.","section":"Section III-B"},{"comment":"The claim 'for the first time' should be checked against reference [11], which already applies LLMs to wireless network design; the authors should clearly state the specific novel element, such as radio map generation or the software-driving agent aspect.","section":"Section I"},{"comment":"The UI description lists items with inconsistent punctuation ('with \"File path\", \"Contact Us\", and \"Help\", 2) prompt...'); please standardize the list formatting.","section":"Fig. 3 and Section III-A"}],"recommendation":"major_revision","confidential_remarks":"The paper is a useful demonstration, but the coverage/SINR comparison is confounded with the built-in ACO optimizer and the autonomy claim outruns the evidence supplied. I would ask the authors to add a proper baseline (manual execution including ACO) and reliability measurements (success rate over multiple prompts or runs), or to scale back the claims to a proof-of-concept demo. The 'first' claim may need checking against reference [11] and related recent work on LLM-based network design."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a workable proof-of-concept that an LLM agent can drive a commercial radio planning tool (RANPLAN) end to end: it downloads an OSM map, creates an environment, runs radio map generation, and invokes the built-in ACO for network optimization via four natural language prompts. That is real engineering. The operation-count reduction (57 manual interactions down to 4 prompts) is a simple, credible count and the platform itself is described with enough detail to be reproduced roughly.\n\nBut the performance comparison in Fig. 7 does not isolate the LLM agent's contribution. The bars labeled \"without LLM\" appear to be the initial unoptimized coverage/SINR, before the ACO step; the \"with LLM\" bars are after ACO. Since ACO is RANPLAN's own optimizer, the coverage gains are expected even if the LLM just launched it. The paper needs a clean baseline where a human manually executes the same ACO workflow, and ideally several runs with variability.\n\nThe bigger gap is the autonomy claim. The agent relies on regular-expression parsing of the LLM's log file to extract tool calls, and there is no error feedback or fallback mechanism; the conclusions explicitly say \"error feedback will be added\" as future work. The two demonstrated scenarios are curated and successful. No failure rate, no prompt variation, no robustness tests. So the claim that this automates planning \"in any locations on the earth\" is not supported by the evidence. The correct interpretation is: the system works in these two cases, under a working assumption that the LLM's output format always matches the parser.\n\nThe novelty is modest—LLM-agent scaffolding (profiles, tools, memory, planning) is standard, and refs [11]–[14] already apply LLMs to network planning. The specific RANPLAN integration and the operation-count reduction are the incremental contributions. The writing is mostly clear, though the \"first\" claim should be softened.\n\nThis paper is for a practitioner audience: someone wanting to see a concrete demo of LLM-as-orchestrator for commercial planning tools, or a venue like IEEE Wireless Communications Letters or a workshop, where a systems demonstration with clear limitations can still be useful. It deserves referee time, but it needs a revised baseline and an honest discussion of reliability before publication.\n\nRecommendation: send it to peer review, but with a request for a clean comparison and a reliability analysis; otherwise it remains a two-example demo.","headline":"A credible but narrow proof-of-concept that an LLM agent can drive RANPLAN end to end; the autonomy claim needs reliability evidence and the performance comparison needs a clean baseline.","tokens_in":6853,"tokens_out":2866,"would_cite":true,"duration_ms":26856,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that an LLM agent can automate radio map generation and wireless network planning by driving commercial planning software with natural-language prompts, cutting manual operations from 57 to 4 while improving coverage and…","keywords":["large language model","LLM agent","radio map generation","wireless network planning","software automation","coverage enhancement","network optimization"],"falsifier":"Run the same four-prompt agent on a wider set of sites, including a dense downtown, a hilly rural area, and a location with incomplete map data, and record how many runs finish without human menu corrections and produce valid radio maps; a substantial failure rate would show the reported reduction in manual operations does not generalize beyond the two demos.","tokens_in":5996,"feed_emoji":"📡","tokens_out":10019,"duration_ms":88876,"temperature":0.7,"pith_summary":"Radio maps and wireless network plans usually come from commercial planning software that requires dozens of manual menu operations per run. This paper proposes an LLM agent—a language model that calls external tools—that takes a natural-language prompt and carries out the whole pipeline: downloading a map, building the outdoor environment, generating a radio map, and optimizing base-station parameters. In tests on an urban campus and a park, the agent reduced manual operations from 57 clicks to 4 prompts and improved the path-loss and SINR coverage statistics of the planned network. The point, if it holds, is that radio map generation and network planning could become fast, scalable, and usable by non-specialists.","feed_headline":"LLM agents cut radio-map planning from 57 clicks to 4 prompts","feed_subtitle":"Natural-language prompts drive planning software, lifting coverage and SINR in two real-world trials.","key_machinery":"The load-bearing mechanism is the LLM agent's tool-use loop: the profile module tells the model who it is and what it may touch, the tool module wraps the planning software's functions in scripts, and the model module sends prompts to a language-model API and returns a log of chosen tools and parameters. A parser extracts those instructions and executes them, while the task-planning functionality orders steps so that radio map generation precedes network optimization and the memory functionality prevents duplicate actions. In one phrase, the LLM agent is a language model whose reasoning is coupled to external software through tool scripts and a profile.","core_discovery":"The paper's central claim is that a single LLM-based agent can automate both radio map generation and wireless network planning at a level that was previously manual. The agent is built from a profile that fixes its role and constraints, a set of tools that wrap planning-software functions into scripted calls, and a model backend that plans and reasons; short- and long-term memory keep each run from repeating work. Driven only by prompts such as \"import map\", \"create environment\", \"generate radio map\", and \"optimize network\", it executes the complete workflow in two real-world scenarios. The reported outcome is a drop from 57 manual operations to 4 prompts and, after automatic cell optimization, better path-loss coverage (from $52.79\\%$ to $83.48\\%$ of the urban scenario meeting the criterion) and better SINR coverage (from $46.94\\%$ to $59.32\\%$).","pith_inferences":["Editorial inference: the reported saving counts manual operations, not wall-clock time, and deployment still requires supervising the agent's runs; the practical shift is from menu-clicking to prompt-level checking.","Editorial inference: the two demonstration sites are not enough to establish robustness, so the natural next test is running the same four-prompt pipeline over many sites and counting unsupervised completions.","Editorial inference: since the agent only orchestrates existing software functions, the approach is backend-agnostic and could absorb better error-feedback and self-correction loops as those mature."],"forward_implications":["Network operators could plan a site by typing a few prompts instead of stepping through dozens of menus, and could re-run the same flow for new sites by changing a location.","Automatic cell optimization, once wrapped as a tool, explains the reported coverage gains, so the agent's value lies in orchestrating existing optimization engines rather than inventing new network algorithms.","The framework's modular profile-tools-model design should transfer to other engineering software whose functions can be wrapped in scripts.","If the pipeline proves reliable at scale, large numbers of radio maps could be generated automatically, supplying training data for AI-based propagation models."],"supporting_citations":[{"why":"Shows that a large language model can be used for wireless network design, the directly related prior result this agent framework extends.","marker":"[11]"},{"why":"Introduces a collaborative LLM-driven automation framework for intent-based 6G network planning, motivating the high-level automation goal.","marker":"[12]"},{"why":"Proposes LLM-assisted wireless network deployment in urban settings, giving a comparable prior use of LLMs in the same problem space.","marker":"[13]"},{"why":"Applies LLM agents to wireless network slicing, demonstrating that agent-style task execution works in wireless networks and inspiring the memory and planning design.","marker":"[14]"}],"fun_headline_variants":["LLM agent turns 57 manual actions into 4 prompts for network planning","Prompt-driven LLM agent lifts urban coverage to 83% in trials","LLM agent automates radio map and network planning in 4 steps","Natural-language planning: LLM cuts 57 manual ops to 4 prompts"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole reduction in manual work rests on the language model reliably turning each prompt into the correct sequence of tool calls and correctly reading the software's outputs, and the paper demonstrates that reliability in only two tested scenarios.","fun_headline_variants_meta":{"raw":{"variants":["LLM agent turns 57 manual actions into 4 prompts for network planning","Prompt-driven LLM agent lifts urban coverage to 83% in trials","LLM agent automates radio map and network planning in 4 steps","Natural-language planning: LLM cuts 57 manual ops to 4 prompts"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00057,"raw_usage":{"total_tokens":2649,"prompt_tokens":850,"completion_tokens":1799,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":466,"completion_tokens_details":{"reasoning_tokens":1718}},"tokens_in":466,"tokens_out":1799,"duration_ms":12367,"temperature":1.0,"reasoning_tokens":1718,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-10T18:25:41.853305+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same four-prompt agent on a wider set of sites, including a dense downtown, a hilly rural area, and a location with incomplete map data, and record how many runs finish without human menu corrections and produce valid radio maps; a substantial failure rate would show the reported reduction in manual operations does not generalize beyond the two demos.","supporting_citations":[{"cited_title":"Large language model-based wireless network design,","cited_arxiv_id":null,"evidence_quote":"Shows that a large language model can be used for wireless network design, the directly related prior result this agent framework extends."},{"cited_title":"Maestro: LLM-driven collaborative automation of intent-based 6G networks,","cited_arxiv_id":null,"evidence_quote":"Introduces a collaborative LLM-driven automation framework for intent-based 6G network planning, motivating the high-level automation goal."},{"cited_title":"Large language models (LLMs) assisted wireless network deployment in urban settings,","cited_arxiv_id":null,"evidence_quote":"Proposes LLM-assisted wireless network deployment in urban settings, giving a comparable prior use of LLMs in the same problem space."}],"review_version":1}