REVIEW 3 major objections 5 minor 18 references
A cross-region cooperative UAV swarm framework can simultaneously serve dynamic ground traffic and sense aerial targets, reaching about 90% communication QoS and reducing the target-localization Cramér-Rao bound by about 45% relative to a n
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 10:54 UTC pith:K32IT7PY
load-bearing objection The QoS side holds up, but the headline 45% CRB reduction rests on an undefined coherent combining gain and unstated phase-error parameters, so the sensing claim is not checkable as written. the 3 major comments →
UAV Swarming for Air-Ground ISAC via Cross-Region Cooperation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central discovery is that cross-region handshaking—periodic exchange of timing, frequency, and phase references among neighboring regions—can be traded against synchronization energy to materially improve cooperative sensing accuracy without hurting communication service. When the handshaking frequency rises, the residual inter-region phase-error variance falls as σ²_ε = σ²_0 + κ/f_H, which raises the coherent combining gain in the Fisher information matrix and lowers the Cramér-Rao bound. The paper shows that a learned policy can pick a good operating point on this trade-off, cutting the CRB by roughly 45% relative to a no-handshaking baseline while keeping the global QoS near 90%. The
What carries the argument
The framework's core is a service-driven Voronoi partition that assigns UAVs to traffic hotspots; an adaptive handshaking mechanism that sets a common inter-region synchronization frequency f_H(t) from per-UAV preferences weighted by marginal CRB contribution; and a region-level MAPPO with centralized training and decentralized execution that maps local observations to continuous actions for mobility, power, and handshaking preference. The handshaking frequency is the pivotal variable: it links the sensing utility (via phase-error variance and coherence factor) to the energy cost, so the learner can balance all three objectives.
Load-bearing premise
The load-bearing premise is that the residual inter-region phase-error variance follows σ²_ε = σ²_floor + κ/f_H(t) with a finite, unstated floor σ²_floor, and that the Fisher information matrix is scaled by an unstated coherent combining gain α_coh(t); the magnitude of the claimed 45% CRB improvement depends on these unspecified quantities.
What would settle it
Recompute the CRB with an explicit formula for α_coh(t) (e.g., derived from the per-region coherence factors η_ℓ = exp(−σ²_ε,ℓ/2)) and with the phase-error floor σ²_floor set to zero; if the CRB reduction over the no-handshaking baseline falls well below 45% under these reasonable settings, the paper's central sensing claim is not supported.
If this is right
- If the claimed gains hold, UAV swarms can adapt to uneven ground traffic and still keep enough angular diversity to sense aerial targets, reducing the need for dedicated sensing platforms.
- The adaptive handshaking mechanism shows that synchronization need not be perfect or maximally frequent; a learned intermediate frequency can achieve most of the benefit of perfect synchronization at lower energy cost.
- The equivalence of the QoS curves for MAPPO and Without Handshaking indicates that handshaking affects sensing but not communication, cleanly decoupling the two functions.
- The learned trajectories deliberately trade off moving toward the target against serving local vehicles, suggesting that multi-objective coordination is learned rather than hand-coded.
Where Pith is reading between the lines
- If the phase-error model were made explicit (giving a closed-form α_coh), the optimal handshaking frequency could be derived analytically from the trade-off between CRB and energy, potentially replacing the learned policy in stationary environments.
- The same region-partition plus handshaking structure could apply to terrestrial base-station cooperation in ISAC networks, where timing and phase alignment among neighboring cells is also a bottleneck.
- The service-driven partition could be made dynamic, reconfiguring region boundaries as traffic hotspots move, an extension the paper lists as future work.
- One testable extension is to verify the 45% CRB reduction against a model where inter-region phase errors are simulated explicitly rather than through the compact variance formula, to ensure the gain is not an artifact of the assumed floor.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a cross-region cooperative framework for UAV-swarm integrated sensing and communication (ISAC), combining service-driven Voronoi regional partitioning, an adaptive inter-region handshaking mechanism to mitigate residual phase errors, and a region-level multi-agent proximal policy optimization (MAPPO) framework under centralized training and decentralized execution. The optimization objective balances ground communication QoS, aerial-target localization CRB, and UAV energy consumption. Simulation results are reported with a communication QoS of approximately 89--90% and a CRB reduction of about 45% compared with the "Without Handshaking" baseline.
Significance. The problem is timely and the overall architecture is plausible: service-driven regional partitioning and handshaking are sensible ways to couple communication load balancing with cooperative sensing, and the MAPPO formulation follows standard CTDE practice. The clearest strength is the explicit formulation of a joint control problem with well-stated constraints, and the qualitative result that learning can coordinate UAV mobility across regions without collapsing all UAVs onto the target. However, the headline sensing gain—the 45% CRB reduction—is not currently verifiable because the coherent combining gain in Eq. (9) is never defined and the phase-error model parameters in Eq. (8) are not reported. The manuscript therefore does not yet provide a reproducible or falsifiable basis for its central quantitative claim.
major comments (3)
- [Section III-B, Eqs. (8)--(9)] The claimed CRB reduction is determined by undefined quantities. Eq. (8) models residual inter-region phase-error variance as σ²_ϵ,ℓ(t) = σ²_ϵ,ℓ,0 + κ_ℓ/f_H(t), but neither σ²_ϵ,ℓ,0 nor κ_ℓ are given anywhere, and no measurement or calibration source is cited. More importantly, Eq. (9) multiplies the entire FIM by α_coh(t), described only as "determined by the inter-region coherence factors," with no explicit formula connecting α_coh to the coherence factors η_ℓ(t) = exp[−σ²_ϵ,ℓ(t)/2]. Since CRB = Tr[J_s^{-1}], any desired CRB reduction can be produced simply by choosing α_coh as a function of f_H(t). As written, the 45% number is a consequence of unspecified modeling choices, not a derived or measured outcome. Please provide an explicit expression for α_coh in terms of the η_ℓ(t) and the sensing geometry, and report the parameter values used for Eq. (8). Without this, the central sensin
- [Section V (Simulations)] The simulation section omits essentially all key parameter values needed for reproduction. No values are reported for the carrier frequency f_c, bandwidth B, noise PSD N_0, antenna gains G_main/G_side, RCS σ_RCS, energy coefficients ζ_H, weights ω_c/ω_s/ω_e, normalization constants η_CRB/E_max, or the phase-error parameters σ²_ϵ,ℓ,0 and κ_ℓ. The MAPPO hyperparameters (learning rates, network sizes, GAE parameter, clipping value, number of seeds) are also absent. Figures 3 and 4 appear to show single-run learning curves without confidence intervals or multiple seeds. Given that the CRB result is particularly sensitive to the under-specified sensing model, the authors should report a complete parameter table, averaged results over multiple random seeds, and, ideally, a sensitivity analysis showing how the claimed 45% reduction varies with plausible choices of σ²_ϵ,ℓ,0, κ_ℓ, and α_coh.
- [Eq. (16) and Section IV-A] The reward function in Eq. (16) directly contains the same Φ_CRB(t) that is later used as the performance metric. Reward design that incorporates the objective is common and not itself an error, but in this paper it creates a circularity concern: the agent is trained to minimize exactly the CRB that Eq. (9) defines, and the reported improvement is then attributed to the handshaking mechanism. Because the sensing model is under-specified, this circularity is not merely cosmetic. I recommend adding an independent evaluation protocol—for example, fixing a pre-trained policy and evaluating CRB under the derived α_coh expression, or comparing against handshaking-frequency sweeps—so that the sensing gain is shown to follow from the mechanism rather than from the reward shaping.
minor comments (5)
- [Abstract vs. Section V] The abstract states a QoS of "approximately 90%," while Section V reports "approximately 89%." Please harmonize these numbers.
- [Section V, Fig. 4] The observation that the Proposed MAPPO and Without Handshaking QoS curves coincide is explained in the text as expected because handshaking does not affect communication QoS. This is fine, but it would be clearer to state the equivalence of the two curves explicitly in the figure legend or caption rather than only in the main text.
- [Notation] The notation fMm for the neighboring-region set is unusual and visually confusable with a function or index set. A calligraphic symbol or an explicit set notation would improve readability.
- [Section III-B, Eq. (7)] The residual phase-error variance in Eq. (7) includes σ²_pn,ℓ(t) for oscillator phase noise, while the following paragraph also lists phase noise as a non-trackable floor. Please clarify whether σ²_pn,ℓ(t) is included in the floor or is itself handshaking-dependent, as this distinction is central to Eq. (8).
- [References and reproducibility] No code or data availability statement is provided. For a learning-based paper whose quantitative claims rest on simulation, a reproducibility statement or public code release would substantially increase confidence. Also, reference [6] contains a typo ("AAV" should be "UAV").
Circularity Check
CRB reduction is largely built into the phase-error/FIM model and the reward; magnitude depends on unstated parameters.
specific steps
-
self definitional
[Section III-B, Eqs. (8)-(10); Section IV-A, Eq. (16); Section V, Fig. 3]
"the residual inter-region phase-error variance is modeled as σ²_ϵ,ℓ(t) = σ²_ϵ,ℓ,0 + κ_ℓ/f_H(t) ... the Fisher information matrix ... is given by J_s(t) = α_coh(t) Σ_{n∈C_m} ι^s_n(t) a_n(t) a_n^T(t) ... Φ_CRB(t) = Tr[(J_s(t))^{-1}] ... In the converged stage, Adaptive Handshaking reduces the CRB by about 45% compared with the Without Handshaking case. This verifies that the adaptive handshaking mechanism can effectively improve cross-region coherent sensing accuracy."
Equation (8) defines the residual inter-region phase-error variance as decreasing with handshaking frequency f_H(t) (the κ_ℓ/f_H term), and Eq. (9) makes the FIM scale with α_coh(t), which the paper says is 'determined by the inter-region coherence factors' (which depend on exp[-σ²_ϵ,ℓ/2]). Eq. (10) inverts that FIM, so the CRB is, by construction, a decreasing function of f_H(t). The reward in Eq. (16) also directly penalizes log(1+Φ_CRB). Consequently, Section V's demonstration that Adaptive Handshaking 'achieves a lower CRB' and its claim that this 'verifies' the mechanism restate the model's assumptions rather than independently test them. The quantitative 45% is not derivable because α_coh(t) is never given an explicit formula and σ²_ϵ,ℓ,0 and κ_ℓ are not specified, so the magnitude i
full rationale
The paper's QoS claim is defined independently by Eqs. (4)-(6) and is not circular. The sensing gain, however, is a different matter: Eq. (8) enshrines the benefit of handshaking, Eq. (9) enshrines the benefit of coherence, and the reward (16) directly optimizes CRB, so the qualitative conclusion that handshaking lowers CRB follows from the paper's own definitions. This is not a full tautology because the 45% magnitude could, in principle, be a nontrivial output of the simulation, but the absence of any formula for α_coh(t) and any values for σ²_ϵ,ℓ,0 and κ_ℓ makes the reported number unfalsifiable and effectively determined by hidden modeling choices. There is no load-bearing self-citation chain: references [1] and [4] by co-author Gao are ordinary background citations, and no uniqueness theorem or ansatz is imported from the authors' prior work. Overall, partial circularity in the central sensing claim, but not complete definitional equivalence, hence score 5.
Axiom & Free-Parameter Ledger
free parameters (6)
- σ²_ϵ,ℓ,0 (residual phase-error floor)
- κ_ℓ (handshaking-sensitive coefficient)
- α_coh(t) (coherent combining gain)
- ζ_H (handshaking energy coefficient)
- ω_c, ω_s, ω_e, η_CRB, E_max
- MAPPO hyperparameters (learning rates, network sizes, GAE, clip, etc.)
axioms (5)
- domain assumption Free-space LoS path loss and directional antenna with main/side lobes (Eqs. (2)-(4))
- domain assumption Positions quasi-static within each time slot
- ad hoc to paper Residual inter-region phase-error variance model σ²_0 + κ_ℓ/f_H(t) (Eq. (8))
- domain assumption Coherence factor η_ℓ=exp(-σ²_ϵ,ℓ/2) and coherent FIM with gain α_coh (Eq. (9))
- domain assumption MAPPO with CTDE converges to a good policy in this non-stationary multi-agent setting
read the original abstract
To serve the volumetric air-ground space, uncrewed aerial vehicles (UAVs) are urgently needed. Yet, relying on them for integrated sensing and communication (ISAC) introduces two key challenges: 1) dynamic and imbalanced ground communication demand, and 2) limited observation diversity for sensing. To address these issues, a cross-region cooperative framework is designed to coordinate UAV swarms. Specifically, a service-driven regional partitioning scheme is proposed to support traffic-aware UAV communication, and an adaptive handshaking mechanism is introduced to improve cooperative sensing accuracy by mitigating residual inter-region phase errors with controlled synchronization overhead. Based on these designs, a region-level multi-agent proximal policy optimization (MAPPO) framework with centralized training and decentralized execution (CTDE) is developed for cross-region cooperative decision-making. Simulation results demonstrate that the proposed method achieves a communication quality-of-service (QoS) of approximately 90% and reduces the Cram\'er-Rao bound (CRB) by about 45% compared to conventional baselines.
Figures
Reference graph
Works this paper leans on
-
[1]
Integrated sensing, communication, and computation for low-altitude networks towards seamless connectivity and connected intelligence,
S. Gaoet al., “Integrated sensing, communication, and computation for low-altitude networks towards seamless connectivity and connected intelligence,”IEEE Internet of Things Magazine, vol. 9, no. 3, pp. 63–71, May 2026
2026
-
[2]
Integrated sensing and communications: Toward dual- functional wireless networks for 6G and beyond,
F. Liuet al., “Integrated sensing and communications: Toward dual- functional wireless networks for 6G and beyond,”IEEE Journal on Selected Areas in Communications, vol. 40, no. 6, pp. 1728–1767, Jun. 2022
2022
-
[3]
Cooperative ISAC-empowered low-altitude economy,
J. Tanget al., “Cooperative ISAC-empowered low-altitude economy,” IEEE Transactions on Wireless Communications, vol. 24, no. 5, pp. 3837–3853, May 2025
2025
-
[4]
Synesthesia of machines (SoM)-enhanced ISAC precoding for vehicular networks with double dynamics,
Z. Yang, S. Gao, X. Cheng, and L. Yang, “Synesthesia of machines (SoM)-enhanced ISAC precoding for vehicular networks with double dynamics,”IEEE Transactions on Communications, vol. 73, no. 9, pp. 7967–7984, Sept. 2025
2025
-
[5]
UA V-assisted NOMA for enhancing ISAC: A deep reinforcement learning solution,
A. Amhaz, M. Elhattab, S. Sharafeddine, and C. Assi, “UA V-assisted NOMA for enhancing ISAC: A deep reinforcement learning solution,” IEEE Communications Letters, vol. 29, no. 2, pp. 249–253, Feb. 2025
2025
-
[6]
Multiagent deep reinforcement learning for AA V- RIS-assisted integrated sensing and communication,
A. M. Huroonet al., “Multiagent deep reinforcement learning for AA V- RIS-assisted integrated sensing and communication,”IEEE Internet of Things Journal, vol. 12, no. 19, pp. 40083–40097, Oct. 2025
2025
-
[7]
Deep reinforce- ment learning based resource allocation and trajectory planning in inte- grated sensing and communications UA V network,
Y . Qin, Z. Zhang, X. Li, W. Huangfu, and H. Zhang, “Deep reinforce- ment learning based resource allocation and trajectory planning in inte- grated sensing and communications UA V network,”IEEE Transactions on Wireless Communications, vol. 22, no. 11, pp. 8158–8169, Nov. 2023
2023
-
[8]
A joint UA V trajectory, user association, and beam- forming design strategy for multi-UA V-assisted ISAC systems,
R. Zhanget al., “A joint UA V trajectory, user association, and beam- forming design strategy for multi-UA V-assisted ISAC systems,”IEEE Internet of Things Journal, vol. 11, no. 18, pp. 29360–29374, Sept. 2024
2024
-
[9]
Cooperative trajectory planning and resource allocation for UA V-enabled integrated sensing and communication systems,
Y . Panet al., “Cooperative trajectory planning and resource allocation for UA V-enabled integrated sensing and communication systems,”IEEE Transactions on V ehicular Technology, vol. 73, no. 5, pp. 6502–6516, May 2024
2024
-
[10]
MARL-based UA V trajectory and beamforming optimization for ISAC system,
Q. Gao, R. Zhong, H. Shin, and Y . Liu, “MARL-based UA V trajectory and beamforming optimization for ISAC system,”IEEE Internet of Things Journal, vol. 11, no. 24, pp. 40492–40505, Dec. 2024
2024
-
[11]
Hierarchical task offloading for vehicular fog computing based on multi-agent deep rein- forcement learning,
Y . Hou, Z. Wei, R. Zhang, X. Cheng, and L. Yang, “Hierarchical task offloading for vehicular fog computing based on multi-agent deep rein- forcement learning,”IEEE Transactions on Wireless Communications, vol. 23, no. 4, pp. 3074–3085, Apr. 2024
2024
-
[12]
Cluster-based multi- agent task scheduling for space–air–ground integrated networks,
Z. Wang, G. Sun, Y . Wang, H. Yu, and D. Niyato, “Cluster-based multi- agent task scheduling for space–air–ground integrated networks,”IEEE Transactions on Cognitive Communications and Networking, vol. 12, pp. 29–42, Mar. 2025
2025
-
[13]
Hierarchical multi-agent DRL-based dynamic cluster reconfiguration for UA V mobility management,
I. A. Meeret al., “Hierarchical multi-agent DRL-based dynamic cluster reconfiguration for UA V mobility management,”IEEE Transactions on Cognitive Communications and Networking, vol. 12, pp. 4957–4971, Dec. 2025
2025
-
[14]
ISAC-enabled multi- UA V cooperative perception and trajectory optimization,
Q. Wang, R. Chai, R. Sun, R. Pu, and Q. Chen, “ISAC-enabled multi- UA V cooperative perception and trajectory optimization,”IEEE Internet of Things Journal, vol. 11, no. 24, pp. 40982–40995, Dec. 2024
2024
-
[15]
Cooperative ISAC networks: Performance analysis, scaling laws, and optimization,
K. Meng, C. Masouros, A. P. Petropulu, and L. Hanzo, “Cooperative ISAC networks: Performance analysis, scaling laws, and optimization,” IEEE Transactions on Wireless Communications, vol. 24, no. 2, pp. 877– 892, Feb. 2025
2025
-
[16]
Network-level inte- grated sensing and communication: Interference management and BS coordination using stochastic geometry,
K. Meng, C. Masouros, G. Chen, and F. Liu, “Network-level inte- grated sensing and communication: Interference management and BS coordination using stochastic geometry,”IEEE Transactions on Wireless Communications, vol. 23, no. 12, pp. 19365–19381, Dec. 2024
2024
-
[17]
Cooperative sensing for ISAC: Challenges, system design, beam management, and performance validation,
G. Liuet al., “Cooperative sensing for ISAC: Challenges, system design, beam management, and performance validation,”IEEE Journal on Selected Areas in Communications, vol. 44, pp. 608–625, Sept. 2025
2025
-
[18]
A novel inte- grated sensing and communication scheme in UA Vs-enabled vehicular networks with MARL-driven adaptive control,
Z. Wang, X.-P. Zhang, W. Ding, Y . Dong, and X. Chen, “A novel inte- grated sensing and communication scheme in UA Vs-enabled vehicular networks with MARL-driven adaptive control,”IEEE Transactions on Mobile Computing, vol. 25, no. 1, pp. 132–147, Jan. 2025
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.