{"id":"c6d4b95c-6a64-4584-96dc-b6fbd82c0557","arxiv_id":"2505.14209","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"The paper derives optimal breach and interception strategies for a 3D perimeter-defense game and introduces an embedded mean-field actor-critic method for large-scale defender coordination.","lead":"This paper studies a 3D perimeter-defense game in which attackers try to breach a protected hemisphere while defenders try to intercept them, and proposes an embedded mean-field actor-critic algorithm that lets many heterogeneous defenders coordinate. It also derives a candidate Nash equilibrium for the one-on-one case and reports simulation and small real-world drone experiments.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Theorem 1's defender-optimality proof contains an algebraic error: Eq. (4) is not a consequence of the stated normalization and is in fact impossible under the triangle inequality. The Fig.","rationale":"I read the paper as claiming two things: a pairwise Nash equilibrium for a one-on-one perimeter-defense game, and a scalable MARL framework (EMFAC) that outperforms baselines. The empirical EMFAC direction is plausible: ablations degrade in the expected directions, runtime overhead is modest, and the real-world 2v2 results are consistent with the simulation story, though no code is released and there is no formal verification. The theoretical part is the weak spot. The reader identified the restriction to deviations C on AB* as a key concern, but that restriction is actually not the core problem: because the attacker's path is fixed to AB*, any non-boundary interception must occur on that segment, and a simple triangle-inequality argument (P(C)-P(B*) = ||DC||-||DB*||+c/v ≥ 0) shows no such C improves the defender's scalar payoff. The real issue is that the proof as written asserts Eq. (4), which is algebraically false under the stated normalization and reduces to a triangle-inequality violation. This is a genuine flaw in the manuscript's central proof, even though the theorem's conclusion appears to be salvageable. An additional, independent mistake is that the simulation in Section V.A.1 reverses the speeds: the theory assumes defender speed 1 and attacker speed v≤1, but the experiment uses attacker speed 1.0 and defender speed 0.8. Thus what is presented as validation of the Nash equilibrium does not actually test the derived equilibrium. These problems justify the reader's CONDITIONAL verdict: the theoretical claim needs repair and re-verification, but the empirical contribution is not convincingly refuted.","tokens_in":19151,"tokens_out":22033,"duration_ms":221192,"concrete_test":"Recompute the defender-optimality step symbolically for a generic configuration: set L=||AB*||, c=||CB*||, and check whether Eq. (4) follows from ||A'B*||=L/v; if it reduces to ||DC||+c<||DB*||, the proof is invalid. Then verify the corrected inequality P(C)-P(B*)=||DC||-||DB*||+c/v ≥ 0 via the triangle inequality, and rerun the Fig. 4 experiment with the model's actual speeds (defender speed 1, attacker speed v=0.8) instead of the reversed speeds used in the paper.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central theoretical claim is the Nash equilibrium in Theorem 1. The defender-optimality step defines A' by ||A'B*||=||AB*||/v and asserts Eq. (4) for any C on AB*. Writing L=||AB*||, c=||CB*||, and d=||DB*||, with C between A and B*, the normalized distance is ||A'C||=L/v - c, whereas the attacker's true travel time to C is (L-c)/v; these differ by c(1/v-1). Substituting ||A'C||=L/v-c into Eq. (4) reduces it to ||DC|| + c < d. This is never true because the triangle inequality gives d ≤ ||DC|| + c. The printed proof therefore does not establish the claimed no-deviation property. The claim is repairable: P(C)-P(B*) = ||DC|| - d + c/v ≥ -c + c/v = c(1/v - 1) ≥ 0 for v ≤ 1, so no point on AB* improves the defender's scalar payoff, and any interception must occur on AB* since the attacker's trajectory is fixed. Separately, Section V.A.1 sets attacker speed 1.0 and defender speed 0.8, the reverse of the model's defender-unit-speed and attacker-speed-v≤1 assumption, so the simulation labeled as Nash-equilibrium verification does not actually test the derived equilibrium. These gaps do not refute the empirical MARL results but leave the theoretical support overstated.","agreement_with_reader":"partial"},"referee_report":null,"author_rebuttal":null,"desk_editor":null,"rs_alignment":null,"lean_confirmation":null,"pith_extraction":null,"created_at":"2026-08-07T15:38:49.039029+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":null,"supporting_citations":[],"review_version":1}