{"id":"ffd808d5-50f8-4351-8968-3bbe8eadc67e","arxiv_id":"2608.09349","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A perspective paper framing Software-Defined Defence as a three-dimensional engineering challenge and proposing a continuous civilian-to-defence engineering loop, with a 2026-2030 roadmap.","lead":"This white paper argues that defence software should be built through a continuous, DevOps-style engineering loop that transfers proven civilian tools into adversarial military settings. It proposes a 2026-2030 roadmap for validating model-based design, simulation-based testing, resilient connectivity, and edge AI under defence conditions.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The central premise that search-based testing transfers to adversarial defence by retargeting its fitness objective (Section 5.2) is asserted, not demonstrated; an active adversary may break the stationary-input assumption the search mechanism relies on.","rationale":"The reader's weakest_assumption correctly identifies that civilian methods transfer to adversarial defence only if redirecting test objectives and adding hardening is sufficient, without a qualitative change in method. I agree, and I specify the load-bearing point precisely: the domain-agnosticism claim about search-based testing (Section 5.2) rests on a stationary-input assumption that an adaptive adversary violates. This is the single most load-bearing concern because the entire 2026-2030 roadmap is built on using existing methods with redirected objectives; if Simulate & Test requires a fundamentally game-theoretic approach, the timeline and the claim of sustainability collapse. The proposed test directly probes this by checking whether the unmodified pipeline can discover known adversarial failure modes. The paper's honesty about open gaps like TSN-to-DTN is a credit, but it does not mitigate the unsupported assertion in the central transferability premise. The reader's CONDITIONAL verdict remains appropriate, and my analysis does not change it.","tokens_in":27783,"tokens_out":4333,"duration_ms":45030,"concrete_test":"Run the OpenSBT framework (as in [43]) unmodified on a representative UAS navigation model (e.g., PX4 in Gazebo) with GPS-denied/spoofing as the fitness objective, and compare the discovered failure scenarios against a known GPS-spoofing attack suite from at least two independent published studies. If the pipeline fails to recover known attack classes, or only finds scenarios explicitly parameterized in the simulator, then the 'domain-agnostic' claim in Section 5.2 is falsified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's roadmap (Sections 5.2 and 6.1) depends on the claim that OpenSBT's search mechanism is 'domain-agnostic' and that applying it to a UAS navigation model with jamming or spoofing as the fitness objective 'could surface' failure modes before fielding. This assumes adversarial activities can be modeled as fixed input-space parameters drawn from a stationary distribution. In reality, an adversary observes the system and adapts, making the search space a game rather than a scenario-parameter space. The same pipeline may only find failures that the modeler hard-coded as parameters, missing adaptive strategies such as reactive GPS spoofing or communications jamming that respond to the system's state. This is not merely a validation gap; it is a potential qualitative break in the transfer, and it is the kind of assumption the paper asserts rather than demonstrates. The paper itself hints at this strain in Section 5.3, where the TSN-to-DTN handoff is admitted to be 'unsolved in both the civilian and defence literature', but it does not revisit whether the other capabilities, especially Simulate & Test, require more than redirection.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This white paper argues that the central engineering challenge of Software-Defined Defence (SDD) is the 'lifecycle paradox'—the mismatch between decade-long hardware procurement cycles and software and AI that must evolve on a timescale of days—and proposes a continuous DevOps-style engineering loop built from four core capabilities (model-based systems engineering, simulation-based testing, tactical connectivity, low-power edge execution) plus three cross-cutting concerns (cybersecurity compliance, variability management, continuous assurance). For each capability, the paper assesses civil-domain maturity, often using fortiss's own tools and projects as illustrative evidence, assigns TRL levels, identifies open transfer problems, and offers a 2026–2030 roadmap for validation under adversarial and defence-certified conditions. The paper is written as a perspective and a call to action for researchers, industry, policymakers, and defence agencies.","tokens_in":27958,"tokens_out":6260,"duration_ms":61824,"significance":"If the central transfer premise holds, the paper provides a valuable and actionable framing for defence digitalization: it names the lifecycle paradox as the core problem, connects it to concrete engineering capabilities, and offers a time-horizoned roadmap. A notable strength is the paper's candour about open gaps—most prominently the TSN-to-DTN handoff and the lack of adversarial validation for search-based testing—which makes the assessment more credible than typical white papers. The paper also points to open-source tools and standards (AutoFOCUS3, OpenSBT, IETF DetNet/RAW), which are checkable. However, the evidence base is largely self-referential: the maturity assessment for each capability rests on fortiss's own tools and self-assigned TRL levels, and the sustainability claim in the Executive Summary is broader than the evidence supports. Disagreement with consensus is not at issue; rather, the paper's central claim would be strengthened by tempering its scope or providing independent validation.","major_comments":[{"comment":"The claim that the proposed DevOps SDD loop is 'sustainable, considering existing capabilities and methodologies that have been built and proven' in civilian domains is broader than the evidence presented in Section 5. The maturity evidence consists of fortiss's own tools at self-assessed TRL 5-6, none of which has been validated under adversarial or defence-certified conditions, and Section 5.3 explicitly states that the TSN-to-DTN handoff is unsolved. A TRL of 5-6 in benign civilian settings does not by itself establish that the capability will be sustainable when an adversary actively tries to break it. Please soften the sustainability claim to a conditional one, or provide a more detailed maturity assessment with independent evidence and a transparent TRL rubric.","section":"Executive Summary; Section 4"},{"comment":"The transfer claim for search-based testing is asserted rather than demonstrated. The paper states that OpenSBT's 'search mechanism itself is domain-agnostic' and that redirecting the fitness objective to jamming or spoofing 'could surface' failure modes. This assumes that adversarial behaviour can be represented as fixed input-space parameters drawn from a stationary distribution. In practice, an adversary observes the system and adapts, so the search space becomes a game rather than a parameter space; the same pipeline may only find failures that the modeler has explicitly parameterized, missing reactive strategies such as state-dependent GPS spoofing or communications jamming. The paper should either provide evidence that search-based testing can cope with adaptive adversaries—for example, a study applying OpenSBT with an adaptive jamming model—or explicitly state this limitation and explain how the roadmap would address it.","section":"Section 5.2, Section 6.1"},{"comment":"The paper admits that the TSN-to-DTN handoff is 'unsolved in both the civilian and defence literature' (Section 5.3), yet the short-term roadmap (Section 6.1) proposes a 2026–2027 pilot that would 'prove the handoff under representative constraints.' This timeline is difficult to reconcile with the acknowledged research gap. Please clarify whether the pilot is intended to solve an open research problem or to demonstrate an existing partial solution, and specify what evidence would constitute success for the pilot.","section":"Section 5.3, Section 6.1"},{"comment":"The TRL values assigned to each capability are self-assessments by the authors, but the paper does not describe the method by which these levels were determined or whether they were reviewed by independent domain experts. Since the paper's central argument depends on the claimed maturity of these capabilities, the TRL assessment should be presented as a provisional self-assessment rather than as an established classification, or a separate validation process should be described.","section":"Section 5.1-5.5"}],"minor_comments":[{"comment":"The sentence 'Redirected from braking-distance objectives to a UAS navigation or target-identification model, with jamming or spoofing as the fitness objective instead of collision distance, the same pipeline could surface degraded-GPS or adversarial-perturbation failure modes in simulation before fielding' appears verbatim twice in Section 5.2, once in 'Illustrative evidence' and once in 'Maturity and gap.' Remove the duplicate.","section":"Section 5.2"},{"comment":"The entry for 'SyNAPSESystems of Neuromorphic Adaptive Plastic Scalable Electronics' is malformed; it should read 'SyNAPSE (Systems of Neuromorphic Adaptive Plastic Scalable Electronics).'","section":"List of Acronyms"},{"comment":"The text labels 'Mechanical/Structural platform' and 'Electrical/electronical platform' in Figure 2 appear misaligned with the shaded bars; please check the figure's readability and correct the typo 'electronical.'","section":"Section 2.5, Figure 2"},{"comment":"The paper would benefit from a brief statement of its scope as a perspective white paper, distinguishing it from a systematic review or an empirical study, so readers know what evidence standard to expect.","section":"Section 1"}],"recommendation":"major_revision","confidential_remarks":"The manuscript is a white paper and a call to action rather than a conventional research contribution. Its main value is as a framing and roadmap, not as new empirical evidence. The requested revisions—tempering the sustainability claim and clarifying the search-based testing transfer assumption—are intended to bring the paper's claims in line with its evidence base. If the authors address these points, the paper could be suitable for publication as a perspective piece in a software or systems engineering venue."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The thing to know about this paper is simple: it is a well-crafted white paper, not a research result. It names the lifecycle paradox — decade-scale hardware procurement against day-scale software and AI updates — and organises SDD into three engineering dimensions (SSE, AIE, CIE) around a continuous DevOps-style loop. That framing is genuinely useful, even if each piece already exists in the literature it cites. What earns the paper credit is how explicit it is about gaps. It tells you, capability by capability, what is civilian-validated, what TRL it claims, and what is still open — the TSN-to-DTN handoff is admitted unsolved in both literatures, and adversarial validation of the search-based testing is listed as a next step, not a finished job. That honesty is rare in roadmap papers and should be acknowledged.\n\nThe soft spot is the load-bearing transferability premise. The paper asserts, in Section 5.2, that the search mechanism is domain-agnostic and that redirecting its fitness objective to jamming or spoofing is enough to test autonomy under adversarial conditions. The stress-test note is right: an adversary adapts, and the input distribution is no longer stationary. That turns scenario search into a game, and the paper does not demonstrate that the same pipeline can handle it. It hints at the strain in Section 5.3 but does not go back to ask whether Simulate & Test needs more than retargeting. This is a genuine intellectual gap, not a minor omission, but the paper at least creates the right place for the question to be asked.\n\nTwo smaller issues. The maturity evidence is dominated by fortiss's own tools — AutoFOCUS3, OpenSBT, TSNWiFi, DABBER, the neuromorphic lab — which makes the self-assigned TRL ratings a somewhat circular picture. And the Executive Summary's sustainability claim is broader than what the body supports: the evidence says these methods are promising civilian baselines, not that the loop is sustainable under contested conditions. Both are fixable with language changes, but they should be fixed before the roadmap is used to steer funding.\n\nWho is this for? Policymakers, defence agencies, and researchers who want an orientation to SDD engineering challenges and a concrete three-horizon plan. It is not a technical breakthrough, but it is a serious roadmap. I would send it to peer review — a systems-engineering or defence-digitalisation venue — with the expectation of conditional acceptance: temper the sustainability claim, acknowledge the adversarial-adaptation limitation explicitly, and either show or clearly scope the transferability of search-based testing. A serious referee should engage with it.","headline":"Useful and honest SDD roadmap; the transferability premise is its biggest open question, but the paper names that question better than most position papers.","tokens_in":28531,"tokens_out":1659,"would_cite":false,"duration_ms":19852,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This white paper argues that the lifecycle paradox — defence hardware procured over decades while the software on it must change within days — is the central problem Software-Defined Defence must solve, and that a continuous DevOps-style…","keywords":["Software-Defined Defence","lifecycle paradox","continuous engineering loop","model-based systems engineering","simulation-based testing","tactical connectivity","low-power edge execution","civilian-to-defence technology transfer"],"falsifier":"Take the paper's open-source search-based testing framework, redirect its fitness objective from braking distance to sensor denial or jamming, and run it against a public perception or navigation model; then expose the same model to physical jamming in a denied, disrupted, intermittent, and limited testbed. If the simulated failure catalogue misses failures the physical run produces, the claim that the search mechanism is domain-agnostic and that only adversarial validation is missing would be falsified; if the catalogues match, the transfer claim survives.","tokens_in":27543,"feed_emoji":"🔁","tokens_out":14767,"duration_ms":127857,"temperature":0.7,"pith_summary":"This white paper argues that the central engineering problem facing Software-Defined Defence is the lifecycle paradox: defence platforms are procured on decade-long timescales while the software and AI models they carry must be updated within days or hours. Its thesis is that a sustainable solution already exists in civilian engineering practice, provided it is reorganised into a continuous DevOps-style loop that front-loads design and verification, carries them into contested operation, and runs compliance and assurance continuously beneath both. The paper assesses four core capabilities — model-based systems engineering, simulation-based testing, tactical connectivity, and low-power edge execution — and claims each is already proven in automotive, manufacturing, space, or energy settings, with the remaining work being redirection toward adversarial conditions rather than invention from scratch. If the thesis holds, defence procurement and certification can move from one-shot delivery to incremental, evidence-based updates, and the same technology base applies to civilian critical-infrastructure resilience.","feed_headline":"Close defence's decade-versus-days software gap with a DevOps loop","feed_subtitle":"Civilian-proven tools can drive defence software if certification becomes continuous and testing goes adversarial.","key_machinery":"The machinery that carries the argument is the continuous DevOps-style SDD loop together with its central transfer device, the claim that the search mechanism of simulation-based testing is domain-agnostic. The loop has four core stages — model-based systems engineering for design and verification front-loading; simulation-based testing, including automated search-based scenario discovery; tactical and resilient connectivity spanning deterministic (time-sensitive networking), opportunistic, and delay-tolerant (bundle-protocol and named-data networking) regimes; and low-power neuromorphic edge execution — with three cross-cutting concerns, continuous cybersecurity compliance, variability and reuse management, and continuous assurance and resilient AI, running beneath all of them. The transfer device is exemplified by an open-source search-based testing framework that explores an operational design space automatically to expose failure modes: redirected from braking-distance objectives to jamming, spoofing, or degraded-GPS fitness functions, the same pipeline is claimed to surface adversarial failure modes in simulation before fielding. The lifecycle paradox is the framing object that motivates the loop: it is the mismatched-timescale diagram of a platform whose mechanical baseline lasts decades while its AI models are retrained in hours, and the loop is designed to absorb those divergent rhythms.","core_discovery":"On the paper's own terms, the claim is that Software-Defined Defence must be built around a continuous engineering loop rather than a sequential acquisition model. The lifecycle paradox — hardware that lasts decades against software that changes weekly and AI models that change daily — is not a procurement inconvenience but the defining engineering constraint, and it forces three dimensions to be engineered together: software and systems engineering, AI engineering, and connectivity and infrastructure engineering. The paper's constructive claim is that a DevOps-inspired loop can carry a platform from design through contested operation: model-based systems engineering and simulation-based testing front-load verification; tactical connectivity and low-power edge execution carry the design into denied, disrupted, intermittent, and limited conditions; and continuous cybersecurity compliance, continuous assurance, and variability management run beneath both phases. Each of these capabilities, the paper argues, is already validated in civilian and dual-use domains, so the remaining scientific question is not whether the technology is relevant but what must be added, tested, or hardened to make it trustworthy under adversarial conditions — a transfer claim the paper states explicitly for search-based testing, whose search mechanism it describes as domain-agnostic.","pith_inferences":["Applied in reverse, the paper's transfer thesis implies that SDD capability work doubles as civilian critical-infrastructure resilience work: an energy or telecoms network under sustained cyber and disinformation pressure faces the same continuous-update-under-adversary problem, so the same loop applies — a direction the paper notes but leaves undeveloped.","If the domain-agnostic claim holds, the open adversarial-robustness benchmark the paper proposes for 2027–2029 could become a shared alliance-level certification instrument, letting different nations cite the same failure catalogue instead of re-deriving evidence — a consequence the paper leaves implicit.","A testable extension the paper does not run: apply the unmodified search-based testing mechanism to a public perception model with a sensor-denial fitness objective, then compare the simulated failure catalogue with one produced under physical jamming; the size of that simulation-to-reality gap in adversarial conditions is what would decide how far the loop's evidence can be trusted.","The paper's own 2030 horizon may be optimistic: the deterministic-to-delay-tolerant handoff is named as unsolved, and the short-term pilot that would close it is scheduled in the same 2026–2027 window as the adversarial testing meant to feed certification — if the handoff resists standardisation, the medium-term certification pipeline stalls."],"forward_implications":["Defence acquisition and certification can shift from one-shot delivery to incremental, evidence-based updates, where a change to one module triggers re-evaluation of only the affected assurance arguments rather than full platform re-certification.","A system certified in 2025 could carry AI models trained on 2040 data, because continuous assurance and traceability keep a timestamped evidence record that stays current across patches and retraining runs.","Search-based testing redirected at adversarial objectives could produce, in simulation and before fielding, the failure catalogue that jamming, spoofing, and sensor-denial certification currently lack, at a fraction of field-trial cost.","Tactical networks that combine deterministic networking with store-carry-forward and in-network caching could keep mission data flowing across deterministic, opportunistic, and delay-tolerant regimes, provided the deterministic-to-delay-tolerant handoff is closed.","Neuromorphic, event-driven edge execution could keep a soldier-worn or platform-mounted system learning and adapting without a live backhaul link, which is exactly the condition contested environments impose."],"supporting_citations":[{"why":"Formalises Software-Defined Defence and its three dimensions; supplies the definition and the lifecycle-paradox framing the paper builds on.","marker":"[11]"},{"why":"Provides the cyber-physical-systems research agenda whose structural characteristics the SDD argument extends.","marker":"[4]"},{"why":"Supplies the open-source search-based testing framework whose domain-agnostic search mechanism carries the civilian-to-defence transfer claim.","marker":"[43]"},{"why":"Supports the claim that the simulation-to-reality gap can be narrowed, load-bearing for the evidence that simulation-based testing can certify autonomous behaviour.","marker":"[45]"},{"why":"Defines the bundle-protocol store-carry-forward semantics that form the delay-tolerant regime of the tactical connectivity capability.","marker":"[24]"},{"why":"Provides the named-data networking architecture that the connectivity story combines with bundle-protocol transport for denied and intermittent conditions.","marker":"[25]"},{"why":"Emulation evidence that named-data networking outperforms IP in military communications, supporting the connectivity transfer claim.","marker":"[33]"},{"why":"Supplies the standards-track deterministic-networking controller-plane framework that the deterministic connectivity regime claims to extend.","marker":"[31]"},{"why":"Supports the claim that existing safety standards do not reason about learned components, motivating the AI-engineering dimension's certification challenge.","marker":"[16]"}],"fun_headline_variants":["Defence software needs a DevOps loop, not a decade-long procurement","Bridge the lifecycle gap: continuous engineering for defence","Civilian tech, adversarial testing: the SDD transfer path","From decade hardware to daily software: a DevOps fix for defence"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that engineering methods validated in civilian and dual-use settings become trustworthy under actively adversarial conditions — jamming, spoofing, and sensor denial — by redirecting test objectives and adding hardening, without a qualitative change in the method itself.","fun_headline_variants_meta":{"raw":{"variants":["Defence software needs a DevOps loop, not a decade-long procurement","Bridge the lifecycle gap: continuous engineering for defence","Civilian tech, adversarial testing: the SDD transfer path","From decade hardware to daily software: a DevOps fix for defence"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000595,"raw_usage":{"total_tokens":2835,"prompt_tokens":1042,"completion_tokens":1793,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":658,"completion_tokens_details":{"reasoning_tokens":1724}},"tokens_in":658,"tokens_out":1793,"duration_ms":82935,"temperature":1.0,"reasoning_tokens":1724,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-11T18:55:16.096837+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the paper's open-source search-based testing framework, redirect its fitness objective from braking distance to sensor denial or jamming, and run it against a public perception or navigation model; then expose the same model to physical jamming in a denied, disrupted, intermittent, and limited testbed. If the simulated failure catalogue misses failures the physical run produces, the claim that the search mechanism is domain-agnostic and that only adversarial validation is missing would be falsified; if the catalogues match, the transfer claim survives.","supporting_citations":[{"cited_title":"Positionspapier Software Defined Defence,","cited_arxiv_id":null,"evidence_quote":"Formalises Software-Defined Defence and its three dimensions; supplies the definition and the lifecycle-paradox framing the paper builds on."},{"cited_title":"Bundle Protocol Version 7,","cited_arxiv_id":null,"evidence_quote":"Defines the bundle-protocol store-carry-forward semantics that form the delay-tolerant regime of the tactical connectivity capability."},{"cited_title":"Supporting military communications with named data networking: An emulation analysis,","cited_arxiv_id":null,"evidence_quote":"Emulation evidence that named-data networking outperforms IP in military communications, supporting the connectivity transfer claim."},{"cited_title":"A Framework for the Deterministic Networking (DetNet) Controller Plane,","cited_arxiv_id":null,"evidence_quote":"Supplies the standards-track deterministic-networking controller-plane framework that the deterministic connectivity regime claims to extend."},{"cited_title":"A safety case pattern for systems with machine learning components,","cited_arxiv_id":null,"evidence_quote":"Supports the claim that existing safety standards do not reason about learned components, motivating the AI-engineering dimension's certification challenge."}],"review_version":1}