REVIEW 1 major objections 1 minor 25 references
A soft actor-critic reinforcement learning approach with Dirichlet policy allocates power to improve mobile target tracking in ISAC systems while sustaining communication performance.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
Applies SAC DRL with Dirichlet policy to adaptive power allocation in mobile-target ISAC systems and reports improved tracking in simulations while preserving communication rates.
T0 review reviewed 2026-06-27 challenge →
load-bearing objection Standard SAC application to ISAC power allocation shows simulation gains but leaves generalization to new motion models untested. the 1 major comments →
Deep Reinforcement Learning for Adaptive Power Allocation in ISAC Systems with Mobile Target
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The SAC-based DRL scheme with a Dirichlet policy and carefully designed reward function enhances tracking performance compared to other baselines while sustaining communication performance in the simulated ISAC system with mobile target.
What carries the argument
Soft actor-critic reinforcement learning combined with a Dirichlet policy that generates normalized power allocation actions, steered by a reward function balancing sensing and communication metrics.
Load-bearing premise
The Markov decision process formulation and the hand-designed reward function are assumed to produce a policy that generalizes beyond the specific simulation scenarios and target-motion models used in training.
What would settle it
Deploy the trained policy on target trajectories whose motion statistics differ from the training distribution and measure whether tracking error rises markedly relative to a policy retrained on the new distribution.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper models the power allocation problem in an ISAC system for tracking a mobile target as a Markov decision process and solves it using a soft actor-critic (SAC) deep reinforcement learning method combined with a Dirichlet policy for normalized continuous actions. A reward function is designed to balance tracking and communication performance. Simulation results indicate that the proposed scheme improves tracking performance compared to baselines while maintaining communication rates.
Significance. If the results are robust, the work demonstrates a practical application of DRL for adaptive resource allocation in ISAC systems, leveraging the Dirichlet policy to handle continuous normalized power allocations under random target motion. This could be valuable for systems requiring dynamic trade-offs between sensing and communication.
major comments (1)
- [Simulation Results] The reported improvements are obtained under fixed target-motion statistics and a hand-crafted reward function. No experiments test generalization to different process noise, turn rates, or SNR regimes, which is load-bearing for claiming that the DRL approach provides a superior policy rather than one tuned to the training distribution.
minor comments (1)
- The abstract mentions 'other baselines' but does not specify what they are; this should be clarified in the main text.
Simulated Author's Rebuttal
We thank the referee for the detailed review and constructive feedback on our manuscript. Below we provide a point-by-point response to the major comment.
read point-by-point responses
-
Referee: [Simulation Results] The reported improvements are obtained under fixed target-motion statistics and a hand-crafted reward function. No experiments test generalization to different process noise, turn rates, or SNR regimes, which is load-bearing for claiming that the DRL approach provides a superior policy rather than one tuned to the training distribution.
Authors: We agree that the current simulations use fixed target-motion statistics and a specific reward function, and that no additional experiments vary process noise, turn rates, or SNR regimes. These parameters were selected to represent standard mobile-target ISAC tracking scenarios, and the Dirichlet policy is intended to produce normalized actions in a general manner. However, the referee is correct that broader generalization tests are needed to support claims of a superior policy beyond the training distribution. In the revised manuscript we will add simulation results under varied process noise, turn rates, and SNR conditions to address this point. revision: yes
Circularity Check
No circularity; derivation is self-contained simulation study
full rationale
The paper formulates power allocation as an MDP, applies standard SAC with a Dirichlet policy, and hand-designs a reward to trade off sensing and communication objectives. All performance claims rest on simulation comparisons to baselines under the chosen motion models and reward. No equation reduces a prediction to a fitted input by construction, no self-citation chain bears the central result, and no uniqueness theorem or ansatz is imported from prior author work. This is a normal non-circular application of DRL; the reader's score of 2 is consistent with the absence of any load-bearing circular step.
Axiom & Free-Parameter Ledger
Cite this review
Pith. "Pith review of Deep Reinforcement Learning for Adaptive Power Allocation in ISAC Systems with Mobile Target." pith.science (2026). https://pith.science/paper/VQPXHIJ7
@misc{pith2026260612078,
author = {Pith},
title = {Pith review of: Deep Reinforcement Learning for Adaptive Power Allocation in ISAC Systems with Mobile Target},
year = {2026},
howpublished = {\url{https://pith.science/paper/VQPXHIJ7}},
note = {Machine review of arXiv:2606.12078}
}
read the original abstract
In this paper, we study the power allocation for an integrated sensing and communication (ISAC) system which tracks a mobile target. We first model the problem as a Markov decision process, and then tackle it with a soft actor-critic (SAC) based deep reinforcement learning (DRL) approach. We also combine a Dirichlet policy, which naturally produces normalized continuous actions under random target motion. To exploit different features of sensing and communication operations, we carefully design a reward function such that the system can dynamically control power allocation to conserve resources. The simulation results demonstrate that the proposed scheme enhances tracking performance compared to other baselines while sustaining communication performance.
Figures
Reference graph
Works this paper leans on
-
[1]
Enabling Joint Communication and Radar Sensing in Mobile Networks—A Survey,
J. A. Zhang, M. L. Rahman, K. Wu, X. Huang, Y . J. Guo, S. Chen, and J. Yuan, “Enabling Joint Communication and Radar Sensing in Mobile Networks—A Survey,”IEEE Commun. Surv. Tutor ., vol. 24, pp. 306– 345, First quarter 2022
2022
-
[2]
Interworking of DSRC and Cellular Network Technologies for V2X Communications: A Survey,
K. Abboud, H. A. Omar, and W. Zhuang, “Interworking of DSRC and Cellular Network Technologies for V2X Communications: A Survey,” IEEE Trans. V eh. Technol., vol. 65, pp. 9457–9470, Dec. 2016
2016
-
[3]
Framework for a Perceptive Mobile Network Using Joint Communica- tion and Radar Sensing,
M. L. Rahman, J. A. Zhang, X. Huang, Y . J. Guo, and R. W. Heath, “Framework for a Perceptive Mobile Network Using Joint Communica- tion and Radar Sensing,”IEEE Trans. Aerosp. Electron. Syst., vol. 56, pp. 1926–1941, June 2020
1926
-
[4]
IEEE 802.11ad-Based Radar: An Approach to Joint Vehicular Communication-Radar System,
P. Kumari, J. Choi, N. Gonz ´alez-Prelcic, and R. W. Heath, “IEEE 802.11ad-Based Radar: An Approach to Joint Vehicular Communication-Radar System,”IEEE Trans. V eh. Technol., vol. 67, pp. 3012–3027, Nov. 2018
2018
-
[5]
Interference Can- cellation and Iterative Detection for Orthogonal Time Frequency Space Modulation,
P. Raviteja, K. T. Phan, Y . Hong, and E. Viterbo, “Interference Can- cellation and Iterative Detection for Orthogonal Time Frequency Space Modulation,”IEEE Trans. Wirel. Commun., vol. 17, pp. 6501–6515, Aug. 2018
2018
-
[6]
Gonz ´alez-Prelcic, M
N. Gonz ´alez-Prelcic, M. Furkan Keskin, O. Kaltiokallio, M. Valkama, D. Dardari, X. Shen, Y . Shen, M. Bayraktar, and H. Wymeersch, “The Integrated Sensing and Communication Revolution for 6G: Vision, 6 0 5000 10000 150000 10 20 30 40 50 60 70 SAC with pred. states Proposed SAC with pred. states (moving average) Proposed (moving average) (a) Total reward...
2024
-
[7]
Networked Integrated Sensing and Communications for 6G Wireless Systems,
J. Li, X. Shao, F. Chen, S. Wan, C. Liu, Z. Wei, and D. Wing Kwan Ng, “Networked Integrated Sensing and Communications for 6G Wireless Systems,”IEEE Internet Things J., vol. 11, pp. 29062–29075, May 2024
2024
-
[8]
Beamformer Design and Op- timization for Joint Communication and Full-Duplex Sensing at mm- Waves,
C. B. Barneto, T. Riihonen, S. D. Liyanaarachchi, M. Heino, N. Gonz ´alez-Prelcic, and M. Valkama, “Beamformer Design and Op- timization for Joint Communication and Full-Duplex Sensing at mm- Waves,”IEEE Trans. Commun., vol. 70, pp. 8298–8312, Dec. 2022
2022
-
[9]
Integrated Sensing and Channel Estimation by Exploiting Dual Timescales for Delay-Doppler Alignment Modulation,
Z. Xiao, Y . Zeng, F. Wen, Z. Zhang, and D. W. K. Ng, “Integrated Sensing and Channel Estimation by Exploiting Dual Timescales for Delay-Doppler Alignment Modulation,”IEEE Trans. Wirel. Commun., vol. 24, pp. 415–429, Jan. 2025
2025
-
[10]
Deep Learning-Based Link Configuration for Radar-Aided Multiuser mmWave Vehicle-to-Infrastructure Communication,
A. Graff, Y . Chen, N. Gonz ´alez-Prelcic, and T. Shimizu, “Deep Learning-Based Link Configuration for Radar-Aided Multiuser mmWave Vehicle-to-Infrastructure Communication,”IEEE Trans. V eh. Technol., vol. 72, pp. 7454–7468, Jan 2023
2023
-
[11]
Throughput Maximization for UA V-Enabled Integrated Periodic Sensing and Com- munication,
K. Meng, Q. Wu, S. Ma, W. Chen, K. Wang, and J. Li, “Throughput Maximization for UA V-Enabled Integrated Periodic Sensing and Com- munication,”IEEE Trans. Wirel. Commun., vol. 22, pp. 671–687, Jan. 2023
2023
-
[12]
Asynchronous Protocol Designs for Energy Efficient Mobile Edge Computing Systems,
S. Eom, H. Lee, J. Park, and I. Lee, “Asynchronous Protocol Designs for Energy Efficient Mobile Edge Computing Systems,”IEEE Trans. V eh. Technol., vol. 70, pp. 1013–1018, Jan. 2021
2021
-
[13]
On the Design Details of SS/PBCH, Signal Generation and PRACH in 5G-NR,
A. Chakrapani, “On the Design Details of SS/PBCH, Signal Generation and PRACH in 5G-NR,”IEEE Access, vol. 8, pp. 136617–136637, July 2020
2020
-
[14]
Power allocation of integrated sensing and communication system for the internet of vehicles,
Z. Pu, W. Wang, Z. Lao, Y . Yan, and H. Qin, “Power allocation of integrated sensing and communication system for the internet of vehicles,”IEEE Trans. Green Commun. Netw., vol. 8, pp. 1717–1728, Dec. 2024
2024
-
[15]
Sensing as a Service in 6G Perceptive Networks: A Unified Framework for ISAC Resource Allocation,
F. Dong, F. Liu, Y . Cui, W. Wang, K. Han, and Z. Wang, “Sensing as a Service in 6G Perceptive Networks: A Unified Framework for ISAC Resource Allocation,”IEEE Trans. Wirel. Commun., vol. 22, pp. 3522– 3536, Nov. 2023
2023
-
[16]
EKF-Based Beamforming Design for Joint Beam Tracking and Communication Systems,
M. Li, F. Dong, T. Liu, and F. Liu, “EKF-Based Beamforming Design for Joint Beam Tracking and Communication Systems,”IEEE Wirel. Commun. Lett., Early Access, 2025
2025
-
[17]
Intelligent Resource Allocation in Joint Radar-Communication With Graph Neural Networks,
J. Lee, Y . Cheng, D. Niyato, Y . L. Guan, and D. G. Gonz´alez, “Intelligent Resource Allocation in Joint Radar-Communication With Graph Neural Networks,”IEEE Trans. V eh. Technol., vol. 71, pp. 11120–11135, Oct. 2022
2022
-
[18]
Intelli- gent Beam Tracking in Radar-Assisted MIMO-OFDM Communication Systems,
Y . Wang, M. Lou, W. Qian, Y . Bai, L. Tang, and Y .-C. Liang, “Intelli- gent Beam Tracking in Radar-Assisted MIMO-OFDM Communication Systems,”IEEE Trans. V eh. Technol., vol. 73, pp. 16774–16789, June 2024
2024
-
[19]
Cooperative Multiagent Deep Reinforcement Learning Methods for UA V-Aided Mobile Edge Computing Networks,
M. Kim, H. Lee, S. Hwang, M. Debbah, and I. Lee, “Cooperative Multiagent Deep Reinforcement Learning Methods for UA V-Aided Mobile Edge Computing Networks,”IEEE Internet Things J., vol. 11, pp. 38040–38053, Dec. 2024
2024
-
[20]
Multiagent Deep Reinforce- ment Learning for Decentralized Multi-UA V Mobile Edge Computing Networks,
S. Hwang, H. Lee, M. Kim, and I. Lee, “Multiagent Deep Reinforce- ment Learning for Decentralized Multi-UA V Mobile Edge Computing Networks,”IEEE Internet Things J., vol. 12, pp. 14484–14497, May 2025
2025
-
[21]
Radar-Assisted Predictive Beamforming for Vehicular Links: Communication Served by Sensing,
F. Liu, W. Yuan, C. Masouros, and J. Yuan, “Radar-Assisted Predictive Beamforming for Vehicular Links: Communication Served by Sensing,” IEEE Trans Wirel. Commun., vol. 19, no. 11, pp. 7704–7719, 2020
2020
-
[22]
Scaled Accuracy based Power Allocation for Multi-target Tracking with Colocated MIMO Radars,
Y . Yuan, W. Yi, T. Kirubarajan, and L. Kong, “Scaled Accuracy based Power Allocation for Multi-target Tracking with Colocated MIMO Radars,”Signal Processing, vol. 158, pp. 227–240, Mar. 2019
2019
-
[23]
Posterior Cramer-Rao Bounds for Discrete-time Nonlinear Filtering,
P. Tichavsky, C. Muravchik, and A. Nehorai, “Posterior Cramer-Rao Bounds for Discrete-time Nonlinear Filtering,”IEEE Trans. Signal Process, vol. 46, pp. 1386–1396, May 1998
1998
-
[24]
Soft Actor-Critic Algorithms and Applications
T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V . Kumar, H. Zhu, A. Gupta, P. Abbeel,et al., “Soft actor-critic algorithms and applications,”arXiv preprint arXiv:1812.05905, 2018
work page internal anchor Pith review Pith/arXiv arXiv 2018
-
[25]
A prescriptive Dirichlet power allocation policy with deep reinforcement learning,
Y . Tian, M. Han, C. Kulkarni, and O. Fink, “A prescriptive Dirichlet power allocation policy with deep reinforcement learning,”Reliability Engineering & System Safety, vol. 224, p. 108529, Aug. 2022
2022
This paper was first reviewed by grok-4.3 on June 27, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.