REVIEW 3 major objections 2 minor 19 references
A bi-level DRL framework positions movable antennas and designs symbol-level waveforms to raise the minimum radar SINR in DFRC systems.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · grok-4.3
2026-06-29 11:03 UTC pith:PB6YLZX4
load-bearing objection Incremental bi-level TD3-plus-CCP setup for MA placement and symbol-level waveforms in DFRC, but the simulation gains rest on unexamined TD3 behavior in non-convex space. the 3 major comments →
Movable Antenna Enhanced Dual-Functional Radar-Communication: A Symbol-Level Precoding Approach
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The bi-level optimization framework, with the twin delayed deep deterministic policy gradient algorithm in the outer loop selecting antenna positions and penalty convex-concave procedure together with majorization-minimization in the inner loop regularizing the symbol-level precoder and filters, yields improved minimum radar SINR values and a superior sensing-communication trade-off relative to benchmark schemes in cluttered environments.
What carries the argument
Bi-level optimization in which TD3 searches antenna positions to maximize min radar SINR while an inner CCP/MM loop produces compliant space-time waveforms and receive filters, handling the nonlinear position-to-channel mapping.
Load-bearing premise
The TD3 algorithm can locate antenna positions that meaningfully raise the minimum radar SINR without converging to placements that leave the objective unimproved.
What would settle it
A set of Monte-Carlo trials in which randomly chosen antenna positions produce equal or higher min radar SINR than the TD3-derived positions would show that the placement optimization step adds no value.
If this is right
- The minimum radar SINR across targets increases relative to fixed-antenna designs.
- The sensing-communication performance trade-off curve lies above those of the benchmark schemes.
- The approach supplies a tractable surrogate for an otherwise non-convex joint placement-and-waveform problem.
- The non-linear mapping from positions to channels is navigated sufficiently well for the reported SINR gains to appear.
Where Pith is reading between the lines
- The same outer-loop placement search could be paired with other inner-loop solvers if the waveform constraints change.
- Real-time re-optimization of antenna positions would become feasible once the computational cost of the inner CCP/MM iterations is reduced.
- Extension to scenarios with moving targets would require only that the channel model inside the inner loop be updated at each time step.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a bi-level optimization framework for symbol-level precoding in movable-antenna enhanced dual-functional radar-communication (DFRC) systems. The outer loop uses the twin delayed deep deterministic policy gradient (TD3) algorithm to optimize antenna positions while the inner loop applies penalty convex-concave procedure (CCP) and majorization-minimization (MM) to design space-time waveforms and receive filters, with the goal of maximizing the minimum radar SINR across multiple targets in clutter. Simulations are reported to show improved radar SINR and a better sensing-communication trade-off relative to benchmark schemes.
Significance. If the TD3 outer loop reliably identifies antenna placements that improve the min-SINR objective beyond fixed-position baselines, the framework would offer a practical heuristic for a non-convex joint design problem that arises in MA-DFRC. The combination of DRL with established inner-loop convexification techniques is a reasonable engineering approach, but the absence of convergence analysis or robustness checks on the reported gains limits the result's immediate theoretical or practical impact.
major comments (3)
- [Abstract] Abstract and simulation results section: the headline claim that the proposed method 'significantly improves radar SINR' rests on TD3 successfully navigating the non-linear position-to-channel mapping in a multi-target cluttered environment, yet no convergence guarantees, ablation studies on TD3 hyperparameters, random seeds, or comparisons against exhaustive/convex-relaxation position search are provided; without these, the reported superiority could be an artifact of initialization rather than a robust property of the bi-level scheme.
- [Proposed Method] Problem formulation and proposed method sections: the optimization problem is stated to be intractable due to waveform constraints and the non-linear antenna-position mapping, but the manuscript supplies no analysis of how the TD3 actor-critic updates interact with the inner CCP/MM loop or whether the composite objective remains stable under realistic clutter models; this directly affects whether the claimed min-SINR gains are load-bearing or merely simulation-specific.
- [Simulation Results] Simulation results: no error bars, multiple independent runs, or explicit baseline definitions (e.g., fixed-position arrays, random MA placement, or convex-relaxation alternatives) are described, making it impossible to assess whether the reported sensing-communication trade-off improvements exceed statistical variation or post-hoc tuning effects.
minor comments (2)
- [System Model] Notation for the channel coefficients as functions of antenna positions should be introduced earlier and kept consistent between the system model and the TD3 state representation.
- [Abstract] The abstract mentions 'practical waveform constraints' but does not list them explicitly; a short enumerated list would improve readability.
Simulated Author's Rebuttal
We thank the referee for the constructive feedback highlighting the need for stronger empirical validation and clearer presentation of the bi-level framework. We will revise the manuscript to incorporate multiple runs, error bars, explicit baseline definitions, and additional discussion on empirical stability. However, theoretical convergence analysis for the TD3 outer loop remains outside the scope of this letter.
read point-by-point responses
-
Referee: [Abstract] Abstract and simulation results section: the headline claim that the proposed method 'significantly improves radar SINR' rests on TD3 successfully navigating the non-linear position-to-channel mapping in a multi-target cluttered environment, yet no convergence guarantees, ablation studies on TD3 hyperparameters, random seeds, or comparisons against exhaustive/convex-relaxation position search are provided; without these, the reported superiority could be an artifact of initialization rather than a robust property of the bi-level scheme.
Authors: We agree that robustness evidence can be strengthened. In revision we will add ablation results on TD3 hyperparameters (actor/critic learning rates and exploration noise) and report performance averaged over 5 independent random seeds with different initializations to show consistency of the min-SINR gains. We will also note that exhaustive search over continuous positions is computationally prohibitive and that no convex relaxation for the position subproblem is currently available. Theoretical convergence guarantees for TD3 in this setting are not provided, as they are generally unavailable for DRL heuristics and lie beyond the letter's scope. revision: partial
-
Referee: [Proposed Method] Problem formulation and proposed method sections: the optimization problem is stated to be intractable due to waveform constraints and the non-linear antenna-position mapping, but the manuscript supplies no analysis of how the TD3 actor-critic updates interact with the inner CCP/MM loop or whether the composite objective remains stable under realistic clutter models; this directly affects whether the claimed min-SINR gains are load-bearing or merely simulation-specific.
Authors: The outer TD3 treats the inner CCP/MM solver as a black-box reward evaluator that returns the achieved min-SINR for each candidate position vector. We will insert a short paragraph describing the observed training behavior, including that reward curves remain stable across the simulated clutter realizations without divergence. A rigorous analysis of the composite dynamics is intractable because the inner loop is non-differentiable; such analysis is left for future work. The clutter model follows the standard point-target-plus-clutter formulation used in prior DFRC literature. revision: partial
-
Referee: [Simulation Results] Simulation results: no error bars, multiple independent runs, or explicit baseline definitions (e.g., fixed-position arrays, random MA placement, or convex-relaxation alternatives) are described, making it impossible to assess whether the reported sensing-communication trade-off improvements exceed statistical variation or post-hoc tuning effects.
Authors: We will expand the simulation section to (i) explicitly list all baselines (fixed-position ULA, random MA placement within the feasible region, and a non-MA symbol-level precoding benchmark), (ii) average all curves over 10 independent runs that vary both TD3 random seeds and channel realizations, and (iii) include error bars showing one standard deviation. These additions will allow readers to judge whether the reported trade-off gains exceed statistical variation. revision: yes
- Theoretical convergence guarantees or a complete stability analysis of the TD3 actor-critic updates interacting with the non-differentiable inner CCP/MM loop under general clutter models
Circularity Check
No circularity: bi-level optimization framework is independent of simulation outcomes
full rationale
The paper formulates a bi-level optimization problem to maximize minimum radar SINR by jointly designing waveforms, filters, and MA positions, then solves the intractable non-linear problem via an outer TD3 layer for placement and inner CCP/MM layers for waveforms. This structure is presented as a direct algorithmic response to the stated intractability, with no equations or claims reducing the SINR objective or its solution to a fitted parameter, self-citation chain, or renamed input by construction. Simulation results are reported as empirical outcomes of applying this framework, not as derivations that presuppose the claimed gains. The derivation chain therefore remains self-contained against external benchmarks.
Axiom & Free-Parameter Ledger
axioms (1)
- domain assumption The optimization problem is intractable due to practical waveform constraints and the non-linear mapping from antenna positions to channel coefficients.
read the original abstract
This letter investigates a symbol-level precoder design for movable antenna (MA)-enhanced dual-functional radar-communication (DFRC) systems. To enhance radar sensing capabilities, we formulate an optimization problem aimed at maximizing the minimum radar signal-to-interference-plus-noise ratio (SINR) across multiple targets in a cluttered environment. Our approach jointly designs the space-time transmitted waveforms, receiving filters, and antenna placement. However, the resulting problem is intractable to solve due to practical waveform constraints and the non-linear mapping from antenna positions to the corresponding channel coefficients. To address these challenges, we develop a bi-level optimization framework by leveraging deep reinforcement learning (DRL). Specifically, the twin delayed deep deterministic policy gradient (TD3) algorithm is employed in the outer layer to optimize antenna placement, while penalty convex-concave procedure (CCP) and majorization-minimization (MM) techniques are incorporated in the inner layer for regularizing waveform design. Simulation results demonstrate that the proposed method significantly improves radar SINR and achieves a superior sensing-communication trade-off compared to benchmark schemes.
Figures
Reference graph
Works this paper leans on
-
[1]
A vision of 6G wireless sy stems: Applications, trends, technologies, and open research pro blems,
W. Saad, M. Bennis, and M. Chen, “A vision of 6G wireless sy stems: Applications, trends, technologies, and open research pro blems,” IEEE Network, vol. 34, no. 3, pp. 134–142, 2020
2020
-
[2]
A survey of recent advances in optimization methods for wireless communications,
Y .-F. Liu et al. , “A survey of recent advances in optimization methods for wireless communications,” IEEE J. Sel. Areas Commun. , vol. 42, no. 11, pp. 2992–3031, 2024
2024
-
[3]
Enabling joint communication and radar sensing in mobile networks—a survey,
J. A. Zhang et al. , “Enabling joint communication and radar sensing in mobile networks—a survey,” IEEE Commun. Surv. Tutorials , vol. 24, no. 1, pp. 306–345, 2022
2022
-
[4]
A tutorial on interference exploitation via symbol- level precoding: Overview, state-of-the-art and future di rections,
A. Li et al. , “A tutorial on interference exploitation via symbol- level precoding: Overview, state-of-the-art and future di rections,” IEEE Commun. Surv. Tutorials , vol. 22, no. 2, pp. 796–839, 2020
2020
-
[5]
Joint transmit waveform and passive beamforming design for RIS-aided DFRC systems,
R. Liu et al., “Joint transmit waveform and passive beamforming design for RIS-aided DFRC systems,” IEEE J. Sel. Topics Signal Process. , vol. 16, no. 5, pp. 995–1010, 2022
2022
-
[6]
SLP-based dual-functional waveform design for ISAC systems: A deep learning approach,
P . Jiang et al. , “SLP-based dual-functional waveform design for ISAC systems: A deep learning approach,” IEEE Trans. V eh. Technol., vol. 74, no. 7, pp. 11 105–11 119, 2025
2025
-
[7]
Secure transceiver design for discrete RIS enhanced dual-functional radar-communication: A symbol-level pre coding ap- proach,
R. Y ang et al. , “Secure transceiver design for discrete RIS enhanced dual-functional radar-communication: A symbol-level pre coding ap- proach,” IEEE Wireless Commun. Lett. , vol. 14, no. 4, pp. 1034–1038, 2025
2025
-
[8]
Movable antenna enhanced wir eless sensing via antenna position optimization,
W. Ma, L. Zhu, and R. Zhang, “Movable antenna enhanced wir eless sensing via antenna position optimization,” IEEE Trans. Wireless Com- mun., vol. 23, no. 11, pp. 16 575–16 589, 2024
2024
-
[9]
Movable antennas for wireles s commu- nication: Opportunities and challenges,
L. Zhu, W. Ma, and R. Zhang, “Movable antennas for wireles s commu- nication: Opportunities and challenges,” IEEE Commun. Mag. , vol. 62, no. 6, pp. 114–120, 2024
2024
-
[10]
Movable-antenna en hanced multiuser communication via antenna position optimizatio n,
L. Zhu, W. Ma, B. Ning, and R. Zhang, “Movable-antenna en hanced multiuser communication via antenna position optimizatio n,” IEEE Trans. Wireless Commun. , vol. 23, no. 7, pp. 7214–7229, 2024
2024
-
[11]
Movable antenna for wireless communications: Pro- totyping and experimental results,
Z. Dong et al. , “Movable antenna for wireless communications: Pro- totyping and experimental results,” IEEE Trans. Wireless Commun. , vol. 25, pp. 6586–6599, 2026
2026
-
[12]
Modeling and performance an alysis for movable antenna enabled wireless communications,
L. Zhu, W. Ma, and R. Zhang, “Modeling and performance an alysis for movable antenna enabled wireless communications,” IEEE Trans. Wireless Commun., vol. 23, no. 6, pp. 6234–6250, 2024
2024
-
[13]
Robust transceiver design for RIS enhanced dual- functional radar-communication with movable antenna,
R. Y ang et al. , “Robust transceiver design for RIS enhanced dual- functional radar-communication with movable antenna,” IEEE Trans. V eh. Technol., pp. 1–15, 2026
2026
-
[14]
Joint discrete antenna positioning and beamforming optimization in movable antenna enabled full-duplex ISAC n etworks,
Z. Li et al. , “Joint discrete antenna positioning and beamforming optimization in movable antenna enabled full-duplex ISAC n etworks,” IEEE Trans. Wireless Commun. , vol. 25, pp. 7220–7234, 2026. JOURNAL OF LATEX CLASS FILES, VOL. 18, NO. 9, SEPTEMBER 2020 6
2026
-
[15]
Movable Antenna Empow- ered Covert Dual-Functional Radar-Communication,
R. Y ang, N. Wei, Z. Dong et al. , “Movable Antenna Empow- ered Covert Dual-Functional Radar-Communication,” arXiv e-prints , p. arXiv:2601.14868, Jan. 2026
-
[16]
Meta-reinforcement learning optimization for mov able antenna- aided full-duplex cf-dfrc systems with carrier frequency o ffset,
Y . Xiu, W. Lyu, Y . Li, R. Y ang, P . L. Y eoh, W. Zhang, G. Liu, and N. Wei, “Meta-reinforcement learning optimization for mov able antenna- aided full-duplex cf-dfrc systems with carrier frequency o ffset,” IEEE Transactions on Communications , vol. 74, pp. 5803–5819, 2026
2026
-
[17]
Robust optimization for movable antenna-aided cell-free isac with time synchronization errors,
Y . Xiu, Y . Zhao, R. Y ang, W. Lyu, D. Niyato, D. In Kim, G. Li u, and N. Wei, “Robust optimization for movable antenna-aided cell-free isac with time synchronization errors,” IEEE Transactions on Wireless Communications, vol. 25, pp. 10 082–10 097, 2026
2026
-
[18]
Mova ble antenna enabled isac beamforming design for low-altitude a irborne vehicles,
Y . Xiu, S. Y ang, W. Lyu, P . Lep Y eoh, Y . Li, and Y . Ai, “Mova ble antenna enabled isac beamforming design for low-altitude a irborne vehicles,” IEEE Wireless Communications Letters , vol. 14, no. 5, pp. 1311–1315, 2025
2025
-
[19]
Convex optimization,
S. Boyd, “Convex optimization,” Cambridge UP , Mar. 2004
2004
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.