{"id":"11df3d73-aa7a-4918-a583-781aca4b103f","arxiv_id":"2506.08045","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A narrative survey defines 'agentic UAVs' as drones with perception, cognition, control, and communication layers and catalogs applications and challenges across eight domains.","lead":"This paper describes a new class of smart drones that do more than follow flight paths: they see, reason, remember, and cooperate with other machines. It surveys how such drones are being used in farming, disaster rescue, delivery, security, wildlife tracking, construction, and environmental monitoring.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's central claim that agentic UAVs are an established new class rests on deployment anecdotes and preprints instead of field evidence; this evidentiary gap is the load-bearing weak point.","rationale":"The reader's weakest_assumption correctly identifies that the survey's load-bearing premise is the existence of real agentic UAV capability rather than isolated simulations or proposals, and that many supporting references are preprints or design studies. My stress-test agrees with that reading and sharpens it: the most serious problem is not merely that supporting references are preprints, but that several concrete deployment claims in Section 3 have no citation at all. These uncited anecdotes carry the weight of the 'new class' claim, and their presence makes the evidentiary gap more acute than the reader's summary suggests. However, I do not think this requires a different verdict from CONDITIONAL. A survey can legitimately propose a taxonomy and architecture while noting that the field is emergent; the paper already has substantial limitations sections (Sections 4.1-4.3) that acknowledge technical, regulatory, and reliability barriers. The issue is calibration: Section 3 repeatedly phrases pilot-level or hypothetical scenarios as accomplished deployments, and the conclusion calls agentic UAVs a 'paradigm shift.' If the evidence audit fails, the appropriate fix is to soften these claims and frame the architecture as a synthesis of ongoing research rather than an established operational class. That is a revision, not a rejection. The paper also has other internal inconsistencies worth noting but not central to the main claim: the stated domain count varies between seven and eight in the abstract and Section 1.3, and Section 3.6 ends mid-sentence with the fragment 'security officers can' before abruptly moving to Section 3.7. These are editorial defects that weaken confidence in the survey's craftsmanship, but they are secondary to the evidentiary question. I would therefore keep the reader's CONDITIONAL verdict, with the condition that the authors either supply primary field evidence for the flagship deployment claims or explicitly label them as pilot concepts and research directions. My agreement is 'partial' because the reader's weakest_assumption and mine overlap on the evidence gap, but I identify the additional problem that several pivotal claims cite no source at all, which is a more specific and more actionable defect.","tokens_in":42397,"tokens_out":3592,"duration_ms":38050,"concrete_test":"Create an evidence audit of all specific deployment claims in Sections 3.1-3.8, including vaccine delivery in Rwanda/India, Amazon Prime Air and Zipline fleets, Serengeti and Amazon wildlife tracking, African night anti-poaching operations, EU/US border patrol trials, construction sites in Japan/UAE, and Australian iron ore mines. For each claim, record whether it has any citation; then classify each citation as peer-reviewed field trial with quantitative metrics, preprint or proposal, simulation, or unrelated survey. If fewer than two claims trace to a primary source with operational data such as success rates, error rates, or flight statistics, the paper's central assertion that agentic UAVs are an existing operational class is unsupported and should be reframed as a research agenda rather than a paradigm shift.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that agentic UAVs constitute a new class of autonomous aerial systems distinguished by cognitive capabilities, contextual adaptability, and goal-directed behavior (Section 1.2). For this claim to hold, the surveyed literature must demonstrate real systems with the four-layer integrated architecture described in Section 2.1, not merely proposals or simulations. The paper's own evidence does not support this: numerous specific deployment assertions in Section 3 are either uncited or backed only by 2025 arXiv preprints. For example, Section 3.5 states that agentic UAVs delivered vaccines in Rwanda and India 'with onboard AI rerouting them around weather systems or no-fly zones' but gives no citation. The same section cites Amazon Prime Air and Zipline pilot programs as proof of coordinated fleet delivery without providing field results, error rates, or regulatory approvals. Section 3.7 asserts Serengeti and Amazon deployments for elephant and primate tracking, and night-time anti-poaching operations in African reserves, again without citations. Section 3.6 mentions 'EU and U.S. southern borders' trials for 24/7 autonomous border security with no supporting reference. The definition itself in Section 1.2 leans heavily on self-cited preprints (e.g., ref. 12) and other 2025 preprints (refs. 14, 17, 39). If these flagship deployments cannot be traced to primary field data, the 'paradigm shift' conclusion is premature: the paper documents a research agenda and a proposed architecture, not an empirically established class of systems.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript argues that 'agentic UAVs' constitute a new class of autonomous aerial systems, distinguished from traditional UAVs by cognitive capabilities, contextual adaptability, and goal-directed behavior. It proposes a four-layer architecture (perception, cognition, control, communication) in Section 2.1, compares traditional and agentic UAVs in Tables 1 and 2, surveys eight application domains in Section 3, discusses challenges and solutions in Sections 4 and 5, and concludes with a roadmap in Section 6. The paper is structured as a broad, cross-domain survey rather than an empirical study, and its central contribution is a taxonomy and synthesis of the emerging literature.","tokens_in":42659,"tokens_out":3735,"duration_ms":38321,"significance":"If its central claim were fully supported, the paper would provide a useful organizing framework for an emerging research area, bringing together architectural concepts, application domains, and open challenges in one place. It also offers a plausible four-layer architectural vocabulary and a broad domain matrix that could facilitate cross-domain comparison. However, the significance is currently conditional: the paper's load-bearing assertion that agentic UAVs are an established, fielded class rests on numerous uncited deployment anecdotes and recent arXiv preprints rather than on verifiable primary field data. The paper would be more appropriately positioned as a position/taxonomy paper or as an evidence survey with a documented methodology and verified use cases.","major_comments":[{"comment":"The subsection asserts that 'during the COVID-19 pandemic, agentic UAVs were deployed in Rwanda and India to autonomously deliver vaccines and testing kits to remote clinics, with onboard AI rerouting them around weather systems or no-fly zones,' but no citation supports this claim. The same paragraph states that 'in pilot programs by Amazon Prime Air and Zipline, UAV fleets have demonstrated coordinated package delivery to multiple addresses in a single flight window' without providing field results, error rates, or regulatory approvals. Because the paper's central conclusion that agentic UAVs are a paradigm shift rather than a research agenda depends on such real-world deployments, these assertions either need to be backed by primary sources or explicitly qualified as reported plans and pilots.","section":"Section 3.5, Logistics and Smart Delivery"},{"comment":"Multiple operational claims are presented as established fact without citations. Section 3.6 states that 'this capability is currently under trial in the EU and U.S. southern borders for 24/7 autonomous border security.' Section 3.7 states that 'in the Serengeti and Amazon rainforest, agentic UAVs have been used to autonomously locate and follow elephants, big cats, and primates' and that 'several African reserves' operate anti-poaching UAVs at night that relay coordinates to ranger units 'within seconds.' Section 3.4 similarly refers to pilot programs in London and Tokyo metros without references. These are central examples for the paper's multidomain claim, and they are not traceable to any cited study. The authors should either supply verifiable references for each such deployment or rewrite these passages as research directions and hypothetical scenarios.","section":"Sections 3.6 and 3.7, Security and Wildlife Conservation"},{"comment":"The definition of 'agentic UAVs' in Section 1.2 relies heavily on self-cited and very recent preprints (e.g., refs. 12, 14, 17, 39), and the architectural distinction in Table 2 is qualitative, with many 'agentic' entries citing papers on RL-based resource allocation or task assignment that do not, on their face, demonstrate deployed agentic UAVs with reflective control. The paper needs an explicit operational criterion for what counts as an 'agentic' system in the surveyed literature (e.g., demonstrated closed-loop replanning, memory-based adaptation, or language-grounded mission specification), so that the classification is checkable rather than stipulated. Without such a criterion, the claim that a coherent new class exists is circular.","section":"Sections 1.2 and 2.1, Definition and Architecture"},{"comment":"The paper calls itself a systematic synthesis ('This review aims to fill that gap by systematically examining...') but provides no search strategy, inclusion criteria, database list, year range, or quality assessment for the surveyed works. This is a survey paper whose central value is supposed to come from its coverage and synthesis; the absence of methodology makes the coverage impossible to reproduce or independently verify. The authors should add a methodology subsection (or explicitly reposition the paper as a narrative/position survey and remove the word 'systematically').","section":"Sections 1.1 and 1.3, Survey Methodology"}],"minor_comments":[{"comment":"The abstract says the paper explores 'seven high-impact application domains' but then lists eight (precision agriculture, construction and mining, disaster response, environmental monitoring, infrastructure inspection, logistics, security, and wildlife conservation); Section 3 also contains eight subsections. The count should be fixed to eight, or the abstract domain list should be revised.","section":"Abstract and Section 1.1"},{"comment":"Section 3.6 ends mid-sentence with 'security officers can' and no concluding clause; the subsection is incomplete and must be finished.","section":"Section 3.6"},{"comment":"The text uses inconsistent spacing and capitalization for 'UAVs' (e.g., 'UA Vs', 'UAV', 'UA V'), which should be normalized for readability.","section":"Throughout"},{"comment":"The Figure 4 caption labels do not match the in-text references: the caption labels subfigures (a)-(e) for disaster zones, flood/wildfire, swarm coordination, environmental monitoring, and urban infrastructure, but the text in Sections 3.2 and 3.3 refers to them in a different order (e.g., 'Figure 4b' for disaster response and 'Figure 4c and (Figure 4d)' for environmental monitoring). The caption or the in-text callouts should be aligned.","section":"Figure 4 caption"},{"comment":"There is a duplicated token in 'VLMs such as Flamingo, OpenFlamingo, or GPT-4V GPT-4V'; the duplicate 'GPT-4V' should be removed.","section":"Section 5.2"},{"comment":"The citation list '[173, 173, 174]' repeats reference 173; this appears to be a typographical error.","section":"Section 3.2"}],"recommendation":"major_revision","confidential_remarks":"The manuscript's citation base is heavily weighted toward 2025 arXiv preprints and the authors' own prior taxonomy papers (refs. 12 and 13). If the authors can verify or substantially qualify the uncited deployment claims (sections 3.4-3.7), the paper could become a useful, if still broad, survey. Otherwise, the editor may wish to encourage reframing as a position/vision paper rather than a survey of deployed systems. I would also ask the authors to correct the incomplete Section 3.6 and the abstract domain-count error before any acceptance decision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague, here is my read. The paper is best treated as a broad organizing survey of \"agentic UAVs,\" not as evidence that such systems are deployed at scale. The four-layer stack (perception, cognition, control, communication) is a clean restatement of sense-think-act with modern VLA/VLM dressing, and the comparison tables across eight domains are genuinely handy. Someone new to this area would leave with a good map of the vocabulary and the main research threads: edge AI, multimodal sensing, RL/MARL, federated learning, VLMs, and the standard regulatory and ethical barriers. The challenges section is boilerplate but covers the right ground.\n\nThe soft spot is the one the stress-test flags: the load-bearing claim that agentic UAVs are an established new class rests on deployment anecdotes without citations. Section 3.5 asserts vaccine delivery in Rwanda and India with onboard AI rerouting, with no reference. Amazon Prime Air and Zipline are invoked as proof of coordinated fleet delivery, without field results or error rates. Section 3.7 says Serengeti and Amazon deployments track elephants and primates, again uncited. Section 3.6 claims border trials in the EU and U.S. with no reference. These are exactly the claims that would make the survey more than a taxonomy, and right now they do not survive contact with the evidence. The conclusion that agentic UAVs are a paradigm shift needs to be downgraded to \"a rapidly maturing research area with promising pilots,\" or the deployment claims need primary sources.\n\nThere are also editorial failures: the abstract says seven domains and lists eight; Section 3.6 ends mid-sentence; the text says \"seven to eight\" in multiple places. The equations in Section 2 are illustrative definitions, not derivations, which is fine for a survey but should be labeled as such. The lack of a search protocol and inclusion criteria makes the synthesis hard to reproduce, which matters for a survey claiming comprehensiveness. Self-citation in refs. 12 and 13 is not a problem by itself; the taxonomy does not hinge on those.\n\nWho is this for? New researchers and practitioners who want a compact map of the field and a quick comparison across application domains. It is not a source of verified deployment evidence.\n\nRecommendation: send it to peer review, but with a referee who will insist on fixing the unsupported deployment claims, adding a methodology statement, and copy-editing the truncated sections. With those changes it can be a useful reference. Without them, it is a well-organized but unreliable overview.","headline":"Broad, useful taxonomy of agentic UAVs undermined by unsupported deployment claims and visible copy-editing failures; worth reviewing after major revision.","tokens_in":43194,"tokens_out":2785,"would_cite":false,"duration_ms":32273,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that a coherent class of agentic UAVs is emerging, defined by four tightly coupled layers—perception, cognition, control, and communication—and surveyed across seven major application domains.","keywords":["agentic UAVs","agentic aerial intelligence","autonomous aerial intelligence","AI agent integration","aerial communications","unmanned aerial systems","agentic AI","multi-domain survey"],"falsifier":"A reader could compile the published field-test data for the cited systems and check whether any has logged a complete mission with an unscripted route change, an onboard semantic decision, and successful task completion without operator intervention; if none has, the paper's portrait describes a target rather than a current class.","tokens_in":42192,"feed_emoji":"🚁","tokens_out":5287,"duration_ms":58355,"temperature":0.7,"pith_summary":"This survey tries to establish that drone technology has crossed a line: a new class of \"agentic UAVs\" is emerging that does not simply execute preprogrammed flights but perceives, reasons, plans, and communicates to pursue goals with minimal human oversight. The authors care because this would change what designers, regulators, and users can expect from drones in high-stakes settings such as disaster response, precision agriculture, logistics, and infrastructure inspection. The paper offers a definition, a four-layer architecture, a comparison against traditional UAVs, and a cross-domain synthesis to ground the term \"agentic\" in engineering rather than marketing. If the framing is right, researchers and policymakers gain a shared vocabulary and a shared target for evaluation and certification.","feed_headline":"Agentic UAVs are a distinct new class of aerial agents","feed_subtitle":"Four layers—perception, cognition, control, communication—unify drones across farming, rescue, and logistics.","key_machinery":"The central object is the four-layer agentic stack: perception maps raw sensor input to a semantic representation $o_t = \\Phi(s_t)$; cognition selects an action $a_t = \\pi(g, o_t)$; control converts it into actuation $u_t = \\Gamma(a_t, x_t)$; and communication coordinates with other agents through shared maps and V2X links. The paper uses this stack as the mechanism that distinguishes agentic UAVs from scripted drones and as the lens for comparing seven application domains. Vision-language models and edge AI are the enabling technologies that make the cognition and perception layers practical onboard.","core_discovery":"The paper's central claim is that agentic UAVs form a distinct class of autonomous aerial systems with a common architectural template: perception, cognition, control, and communication layers operating in tight feedback loops, with edge AI and vision-language models as key enablers. It argues that these systems reach context-aware autonomy with minimal human supervision, unlike traditional waypoint-following drones, and it reads seven application domains through this template to show shared patterns and transferable capabilities. The stated contribution is a foundational framework for developing, deploying, and governing these systems.","pith_inferences":["The survey is best read as a design manifesto: much of the field evidence it cites comes from preprints, pilot programs, and concept systems, so the empirical status of \"agentic UAV\" remains open.","A testable extension is an autonomy scorecard that measures each of the four layers separately, such as how often replanning happens without operator input and how often vision-language instructions are grounded incorrectly, which would show whether the bottleneck is perception, reasoning, or regulation.","If the architecture is correct, the same stack could apply to ground robots and maritime vehicles, making \"agentic UAV\" one instance of a general embodied-agent template.","The emphasis on vision-language instruction following suggests that semantic grounding failures, rather than flight mechanics, may become the next practical bottleneck; that prediction could be checked with error logs from field trials."],"forward_implications":["If the four-layer architecture is accepted, individual UAV systems can be evaluated layer by layer, making benchmark design and failure diagnosis more tractable.","Claims of Level 4-5 autonomy would carry a concrete standard: a UAV should be able to replan mid-mission from onboard reasoning without a human in the loop.","Vision-language interfaces would move from research novelty to a core requirement for human-UAV collaboration.","Cross-domain transfer becomes plausible: a perception or planning module validated in one sector, such as disaster search and rescue, could be reused in another, such as wildlife patrol, within the same stack.","Regulators would need new certification paths for learned and adaptive policies, since conventional airworthiness criteria are designed for deterministic systems."],"supporting_citations":[{"why":"Supplies the cognitive-architecture view of embodied agents that grounds the four-layer sense-think-act loop.","marker":"[11]"},{"why":"Gives the taxonomy of AI agents versus agentic AI that underlies the paper's definition of agentic behavior.","marker":"[12]"},{"why":"Provides the multi-agent LLM mission-planning system used as evidence that UAVs can follow language-level goals.","marker":"[14]"},{"why":"Frames UAVs meeting LLMs as the basis for agentic low-altitude mobility and semantic perception.","marker":"[17]"},{"why":"Supports the learned decision-engine and edge-AI component of the cognition and control layers.","marker":"[23]"},{"why":"Anchors the multi-agent embodied AI perspective behind the communication and shared-map layer.","marker":"[39]"},{"why":"Supplies the swarm autonomy and learning-memory capabilities claimed for collaborative agentic systems.","marker":"[42]"},{"why":"Introduces the vision-language model architecture the paper relies on for instruction-following UAVs.","marker":"[64]"},{"why":"Provides a benchmark for embodied agents on UAVs, used to show agentic UAV evaluation is becoming concrete.","marker":"[99]"},{"why":"Supports the disaster-response claims with a deep reinforcement learning search-and-rescue application.","marker":"[171]"}],"fun_headline_variants":["Agentic UAVs: a distinct class with a unified four-layer architecture","Survey identifies one architecture for all agentic UAVs","Agentic UAVs: one template, seven domains","Four-layer architecture defines agentic UAVs","Agentic UAVs: beyond waypoint-following drones"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the cited deployments and pilot programs are genuine demonstrations of agentic UAV capability, not simulations, lab prototypes, or concept proposals.","fun_headline_variants_meta":{"raw":{"variants":["Agentic UAVs: a distinct class with a unified four-layer architecture","Survey identifies one architecture for all agentic UAVs","Agentic UAVs: one template, seven domains","Four-layer architecture defines agentic UAVs","Agentic UAVs: beyond waypoint-following drones"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000787,"raw_usage":{"total_tokens":3438,"prompt_tokens":879,"completion_tokens":2559,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":495,"completion_tokens_details":{"reasoning_tokens":2479}},"tokens_in":495,"tokens_out":2559,"duration_ms":19176,"temperature":1.0,"reasoning_tokens":2479,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T05:44:33.778872+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A reader could compile the published field-test data for the cited systems and check whether any has logged a complete mission with an unscripted route change, an onboard semantic decision, and successful task completion without operator intervention; if none has, the paper's portrait describes a target rather than a current class.","supporting_citations":[{"cited_title":"Bedi: A comprehensive benchmark for evaluating embodied agents on uavs.arXiv preprint arXiv:2505.18229, 2025","cited_arxiv_id":null,"evidence_quote":"Provides a benchmark for embodied agents on UAVs, used to show agentic UAV evaluation is becoming concrete."}],"review_version":1}