{"id":"7f212488-68aa-4350-b858-85c7f9d11047","arxiv_id":"2506.16786","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A systematic literature review organizes UAV dependability threats and techniques and proposes eight open research directions.","lead":"This paper is a systematic review of 458 studies on the dependability of drone-based computing and networking systems. It maps the field into threat categories and mitigation techniques, and lists eight areas for future research.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The survey's central synthesis depends on a reproducible and unbiased 458-paper sample, but undisclosed per-database query tweaks, a stated exclusion of hardware/cyber topics that the taxonomy still includes, and non-reproducible Figure 2/3 counts leave the trend and gap analysis unverifiable.","rationale":"I read the paper in good faith as a systematic mapping study, and its internal taxonomy and Avizienis-based classification are broadly reasonable. The reader's weakest_assumption targets sample representativeness, and the manuscript text supports that as the main risk: per-database query adjustments are admitted, exclusion criteria conflict with included content, and no search logs or coding protocols are provided. Table 1's widely varying acceptance rates make the risk concrete rather than hypothetical. I did not find an internal mathematical contradiction in the technique taxonomy, and the future-directions section is plausible, so the concern is not that the synthesis is demonstrably wrong but that its load-bearing quantitative basis is not independently checkable. A supplementary artifact release would settle most of the concern, so UNCHANGED is appropriate rather than REJECT or ACCEPT.","tokens_in":48114,"tokens_out":1736,"duration_ms":24765,"concrete_test":"Require the authors to release a supplementary PRISMA-style artifact containing: (a) the exact final Boolean search string per database and per search date; (b) the per-database hit, deduplication, title/abstract exclusion, and full-text exclusion counts; (c) the full DOI list of all 458 retained papers with inclusion/exclusion tags; and (d) the coding dictionary mapping keywords to Figure 2 categories. Then independently recompute Figure 1 yearly counts and Figure 2 frequencies from that DOI list. If the recomputed counts differ materially, or if the acceptance-rate discrepancy between ISI and Springer/ScienceDirect changes the stated trends, the quantitative conclusions should be revised or explicitly labeled as indicative rather than systematic.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim is that 458 systematically selected papers support the trend analysis, threat taxonomy, and eight future directions. The load-bearing assumption is that this sample is representative and the quantitative synthesis is auditable. That assumption is currently not secured. Section 3.1 says queries were adjusted per database, but the actual strings are not given; Section 8 concedes this 'could introduce inconsistencies.' Section 3.2 says cyber-security and hardware-level reliability were excluded, yet Section 5 includes System Vulnerabilities, cyber-attacks, Hardware Failure, and Component Aging, so the operational inclusion rule is unclear. Table 1 then shows extreme heterogeneity in per-source acceptance rates (ISI 217/376 = 58%, ScienceDirect 72/582 = 12%, Springer 51/704 = 7%), which can distort any aggregate trend if query or screening thresholds differed by source. The Figure 2/3 keyword counts (e.g., reliability 380 mentions, network disconnection 330 mentions) are presented without a coding dictionary, paper-level tallies, or inter-rater agreement numbers. None of this proves the survey is wrong, but it blocks independent verification of the headline quantitative claims and the evidence base for the eight research gaps. The concern is reproducibility, not fraud.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper reports a systematic literature review of dependability in UAV-based networks and computing systems. The authors describe a protocol for searching five digital libraries, retrieving 1,848 initial records and retaining 458 papers after title/abstract and full-text screening. The paper presents quantitative trend analyses of dependability metrics, system types, threats, techniques, and applications; a threat taxonomy in Section 5; a techniques taxonomy in Section 6 grounded in the Avizienis et al. fault-prevention/tolerance/removal/forecasting framework; and eight future research directions in Section 7. A threats-to-validity discussion appears in Section 8.","tokens_in":48306,"tokens_out":4114,"duration_ms":43722,"significance":"If the underlying synthesis is reliable, the survey provides a useful map of 458 papers, a structured classification of threats and techniques, and a prioritized research agenda. The systematic protocol, the explicit grounding of the techniques taxonomy in a well-known external framework, and the acknowledgment of validity threats are strengths. However, the headline quantitative claims and the evidence base for the eight future directions are only as strong as the reproducibility and consistency of the selection and coding process, and several load-bearing points currently obstruct independent verification.","major_comments":[{"comment":"The taxonomy is internally inconsistent. Section 5 says the taxonomy 'categorizes threats into seven major domains' but then enumerates six: Energy Constraint, Network Issues, Hardware Failure, Software Failure, Environmental Impact, and Operational Failures. The RQ2 answer (1) also claims seven categories and adds 'Communication Breakdowns' as a seventh, yet no section or subsection defines or discusses Communication Breakdowns. This inconsistency undermines the claim of a structured synthesis; please either add the missing domain or revise the count and the RQ2 answer to match the six categories actually presented.","section":"Section 5 and RQ2 answer"},{"comment":"The stated exclusion criteria conflict with included content. Section 3.2 says the authors 'excluded studies primarily focused on cybersecurity, hardware-level reliability, or other secondary literature,' but Section 5.4 covers system vulnerabilities, including exploitation by attackers, channel access attacks, and hijacking, and Section 5.3 covers hardware failures including sensor malfunction, actuator failure, and component aging. Section 5.2 also discusses deliberate channel interference. As written, the operational inclusion rule is unclear and could bias the 458-paper sample. Please clarify how security- and hardware-related literature was actually treated, and either revise the exclusion statement or justify the inclusion of these threat categories within the dependability scope.","section":"Section 3.2 vs. Sections 5.3 and 5.4"},{"comment":"The search is not currently reproducible. Section 3.1 gives a single Boolean search string, but Section 8 concedes that 'minor adjustments were made to tailor the queries for specific databases, which could introduce inconsistencies.' The exact per-database queries are not provided anywhere in the manuscript, so the reported 1,848 initial records and the resulting 458-paper sample cannot be independently reproduced. Please include the full query used for each of the five databases, such as in an appendix or supplementary file.","section":"Section 3.1 and Section 8"},{"comment":"The keyword-frequency counts are presented without the coding protocol needed to audit them. The text reports counts such as reliability (380 mentions) and network disconnection (330 mentions), and Figures 2 and 3 show temporal trends, but there is no coding dictionary, no definition of the counting unit (per paper, per occurrence, per tagged sentence), no paper-level tallies, and no inter-rater agreement measures. These counts are the empirical basis for the trend claims and for RQ1, and they indirectly support the gap analysis in Section 7. Please add a description of the coding procedure, the complete keyword dictionary, and ideally a supplementary data file with per-paper coding results.","section":"Section 4.1.1 and Figures 2 and 3"}],"minor_comments":[{"comment":"The table is titled 'Fault Forecasting Techniques,' but the first column of both data rows reads 'Fault Prevention.' Please correct the label to 'Fault Forecasting.'","section":"Table 6"},{"comment":"The sentence 'Mentions of detection and diagnosis techniques, though lower in frequency not shown in the picture, show gradual growth with single-digit' is unclear; if a category is omitted from Figure 3d, please state this in the figure caption or include the category in the figure.","section":"Section 4.1.2"},{"comment":"The phrase 'This numbers shows' should be 'These numbers show.'","section":"Section 4, text near Figure 1"},{"comment":"Several cells contain duplicate bracketed citations, for example [345] in 'Adverse Weather Conditions' and [109], [86], and [291] in 'Component Aging.' Please deduplicate the citation lists.","section":"Table 2"},{"comment":"The text states that the survey 'do[es] not cover security and privacy-related issues,' which is consistent with the exclusion criterion in Section 3.2 but is difficult to reconcile with the security-related content in Section 5.4; aligning these statements is part of major comment 2 above.","section":"Section 2 and Section 3.2"}],"recommendation":"major_revision","confidential_remarks":"This is a survey paper, and the issues identified are fixable within the manuscript's scope: reconciling the six-versus-seven threat taxonomy count, clarifying the scope regarding security and hardware-level reliability, and providing the exact per-database queries and keyword-coding protocol. Addressing these points would materially increase the paper's value and verifiability. The paper fits the journal's scope. I have no concerns about citation practices beyond noting that the authors' own prior works appear as part of the surveyed literature, which is acceptable for a survey. "},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The paper is a systematic mapping study of 458 papers on dependability of UAV-based networks and computing systems. What is actually new is the integration: previous surveys cover security, energy, edge computing, or AI separately, while this one organizes both networking and computing under the Avizienis dependability taxonomy, with threat categories (energy, network, hardware, software, environmental, operational) and technique categories (fault prevention, tolerance, removal, forecasting). That is a legitimate and useful contribution, and the eight future directions (LLM agents, satellite constellations, B5G/6G, DRL, urban air mobility, precision agriculture, operational safety, maintainability) grow out of the surveyed literature in a sensible way.\n\nThe paper also does several things well. The SLR protocol is clearly described, the inclusion/exclusion criteria are explicit, and Section 8 acknowledges the main threats to validity. The authors cite their own earlier work as part of the corpus, which is fine when the method is systematic. The writing is readable, and the tables give a good quick reference to which papers sit in which category.\n\nThe soft spots are real but not fatal. First, internal inconsistencies: the taxonomy says seven threat domains but lists six; RQ2 adds 'Communication Breakdowns' out of nowhere; the stated scope excludes security and hardware-level reliability while Section 5.4 covers system vulnerabilities and Section 5.3 covers hardware failure and component aging; Table 6 labels 'Fault Forecasting' as 'Fault Prevention' in the Technique column. These are addressable but need fixing. Second, and more important, the quantitative claims are not auditable. Section 3.1 says search queries were adjusted per database, but the actual strings are not given; Section 8 concedes this could introduce inconsistencies. The Figure 2/3 keyword counts have no coding dictionary, no paper-level tallies, and no inter-rater agreement numbers. The per-source acceptance rates in Table 1 (217/376 vs 72/582 vs 51/704) vary so much that aggregate trends could be distorted if screening thresholds differed by source. None of this proves the trends are wrong; it just means the central 'eight research gaps' rest on a sample whose representativeness cannot be independently checked.\n\nWho is this for? Researchers working on UAV reliability or dependability, and anyone wanting a starting map of the literature. It deserves a serious referee. I would send it to peer review with a request for major revision, not desk reject it. The fixes are mechanical and the survey's qualitative synthesis is valuable enough to keep.","headline":"A genuinely useful and mostly sound survey of UAV dependability research, but the trend counts and gap analysis need a reproducibility appendix and a few internal consistency fixes before I'd treat it as authoritative.","tokens_in":48853,"tokens_out":2637,"would_cite":true,"duration_ms":30668,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A systematic literature review of 458 papers from 2015 to 2024 maps the threats and countermeasures that decide whether drone networks and their computing systems can be trusted, and ranks eight gaps for future research.","keywords":["UAV dependability","systematic literature review","network reliability","fault taxonomy","edge computing","fault tolerance","drone networks","research trends"],"falsifier":"Re-run the stated Boolean search across the same five databases and the same 2015-to-July-2024 window with two independent screeners, and have independent coders re-classify a random subsample of the 458 papers using the paper's own category definitions; if inter-rater agreement on threat and technique classes is low, or if adding the explicitly excluded cyber-security and hardware-reliability papers changes the top-ranked threats and techniques, the trend claims and the eight research directions would not be stable.","tokens_in":47882,"feed_emoji":"🛸","tokens_out":11964,"duration_ms":114873,"temperature":0.7,"pith_summary":"This survey tries to give the scattered literature on drone dependability a single, reproducible map. The authors searched five databases for UAV reliability, availability, safety, and fault-tolerance work published from 2015 to mid-2024, kept 458 of 1,848 retrieved papers, and sorted them into a threat taxonomy (energy, network, hardware, software, environmental, and operational failures) and a techniques taxonomy derived from the classic four-part fault framework of prevention, tolerance, removal, and forecasting. The result is a quantified picture of the field: network disconnection and hardware faults lead the threat list, energy-aware design leads the prevention techniques, and reliability and performance lead the measured attributes. The same review exposes what the field neglects - maintainability, safety, correction, and verification that a fix has not introduced a new fault - and from those gaps it proposes eight future research directions.","feed_headline":"458 studies mapped: what actually threatens drone dependability","feed_subtitle":"Reliability work fixates on network loss and battery limits; maintainability and post-fix checks go unstudied.","key_machinery":"The argument is carried by two taxonomies and a review protocol. The threat taxonomy names six failure domains - energy constraint, network issues, hardware failure, software failure, environmental impact, and operational failures - and serves as the lens for categorizing all 458 papers by threat type. The technique taxonomy applies the Avizienis et al. (2004) fault framework, whose four named classes are fault prevention, fault tolerance, fault removal, and fault forecasting, with subcategories such as energy- and resource-aware design, redundant architecture, and probabilistic evaluation. The protocol, a systematic mapping study, supplies the corpus itself: a Boolean search string spanning UAV terms, computing-and-networking domains, and dependability attributes, run over five bibliographic databases, followed by two-stage title-and-abstract screening and keyword-frequency counting that drives the trend and gap claims.","core_discovery":"The survey's central claim is that UAV dependability research, though conducted in isolated subfields, forms a coherent and rapidly growing corpus: annual publication counts rise from three in 2015 to 109 in the first seven months of 2024, and the 458 retained papers can be classified along two axes. The threat axis groups the literature into six named categories - energy constraints, network issues, hardware failure, software failure, environmental impact, and operational failures - with network disconnection the single most frequent concern at 330 keyword mentions. The technique axis, built on the Avizienis et al. (2004) fault framework, sorts methods into fault prevention, fault tolerance, fault removal, and fault forecasting, and shows that prevention-oriented work dominates while fault removal and fault forecasting are underrepresented. From these classifications the authors answer their four research questions and argue that eight areas - LLM/AI agents, satellite constellations, B5G/6G integration, multi-agent deep reinforcement learning, urban air mobility, precision agriculture, operational safety, and maintainability - deserve prioritized future investigation.","pith_inferences":["The stated exclusions of cyber-security and hardware-level reliability conflict with sections that do cover system vulnerabilities and hardware failures; a parallel review that includes security papers would likely shift the top-threat ranking, since many vulnerability studies are security-driven.","Keyword-mention counting can double-count a paper that addresses several threats at once; a paper-level allocation could change category sizes and therefore the gap rankings.","The 'other' threat bucket is large and still growing (peaking at 72 mentions in 2023), which suggests the six named categories may miss dominant real-world failure modes, such as regulatory interference or AI misbehavior, that a finer-grained coding would expose.","The taxonomy is described as having seven categories in the research-question answer but six in the body, with 'communication breakdowns' appearing only in the former; merging the two definitions cleanly would remove a small reproducibility risk for anyone coding papers against this map."],"forward_implications":["The field gets a reproducible baseline: per-source acceptance counts and year-by-year keyword frequencies can anchor any future survey or meta-analysis of UAV dependability.","The technique taxonomy quantifies an imbalance - fault prevention dominates, non-regression verification appears in none of the 458 papers, and correction and fault forecasting are sparse - giving concrete openings for research on runtime adaptation and predictive maintenance.","Underrepresented metrics such as safety, maintainability, fault tolerance, and robustness mark territory where new work would not duplicate existing results.","The eight proposed directions each attach to a measured gap, so the agenda (LLM/AI agents, satellite constellations, B5G/6G, multi-agent deep reinforcement learning, urban air mobility, precision agriculture, operational safety, maintainability) is grounded in the corpus rather than speculative.","Application coverage is concentrated in monitoring and surveillance, so rescue, delivery, and maintenance missions remain comparatively open for dependable-UAV research."],"supporting_citations":[{"why":"Supplies the dependability concept (reliability, availability, maintainability, safety) and the four-class fault framework that organizes the entire techniques taxonomy.","marker":"[23]"},{"why":"Defines the systematic mapping methodology whose protocol - search strings, screening stages, classification - generated the 458-paper corpus.","marker":"[260]"},{"why":"A UAV-assisted IoT survey used to position this survey's integrated networking-and-computing scope against QoS-focused reviews.","marker":"[2]"},{"why":"A computation-offloading survey presented as one of the fragmented, single-aspect reviews this work claims to unify.","marker":"[134]"},{"why":"A security-and-privacy survey that marks the scope boundary: security is excluded here while dependability attributes are kept.","marker":"[226]"},{"why":"A machine-learning and Internet-of-Drones survey used as the related-work baseline for AI-driven UAV techniques.","marker":"[118]"},{"why":"A resource-management survey that exemplifies the isolated treatment of computing aspects that this survey synthesizes.","marker":"[363]"}],"fun_headline_variants":["458 studies on drone dependability: network is the top threat","Drone reliability survey finds maintenance and fixes understudied","UAV dependability: prevention dominates, gaps remain in eight areas","What threatens drone networks? Survey maps 458 papers","Drone dependability research: eight future frontiers identified"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The review assumes that the 458 retained papers - selected with per-database search strings and subjective title-and-abstract screening, a limitation the authors concede in their threats-to-validity section - represent the field faithfully enough that the trend counts and the eight proposed future directions genuinely follow from the literature.","fun_headline_variants_meta":{"raw":{"variants":["458 studies on drone dependability: network is the top threat","Drone reliability survey finds maintenance and fixes understudied","UAV dependability: prevention dominates, gaps remain in eight areas","What threatens drone networks? Survey maps 458 papers","Drone dependability research: eight future frontiers identified"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000372,"raw_usage":{"total_tokens":1987,"prompt_tokens":938,"completion_tokens":1049,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":554,"completion_tokens_details":{"reasoning_tokens":967}},"tokens_in":554,"tokens_out":1049,"duration_ms":13376,"temperature":1.0,"reasoning_tokens":967,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T23:37:34.958214+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the stated Boolean search across the same five databases and the same 2015-to-July-2024 window with two independent screeners, and have independent coders re-classify a random subsample of the 458 papers using the paper's own category definitions; if inter-rater agreement on threat and technique classes is low, or if adding the explicitly excluded cyber-security and hardware-reliability papers changes the top-ranked threats and techniques, the trend claims and the eight research directions would not be stable.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Defines the systematic mapping methodology whose protocol - search strings, screening stages, classification - generated the 458-paper corpus."}],"review_version":1}