Pith. sign in

REVIEW 1 major objections 1 minor 25 references

A soft actor-critic reinforcement learning approach with Dirichlet policy allocates power to improve mobile target tracking in ISAC systems while sustaining communication performance.

Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →

Applies SAC DRL with Dirichlet policy to adaptive power allocation in mobile-target ISAC systems and reports improved tracking in simulations while preserving communication rates.

T0 review reviewed 2026-06-27 challenge →

load-bearing objection Standard SAC application to ISAC power allocation shows simulation gains but leaves generalization to new motion models untested. the 1 major comments →

arxiv 2606.12078 v1 pith:VQPXHIJ7 submitted 2026-06-10 eess.SP cs.SYeess.SY

Deep Reinforcement Learning for Adaptive Power Allocation in ISAC Systems with Mobile Target

classification eess.SP cs.SYeess.SY
keywords ISACpower allocationdeep reinforcement learningmobile target trackingsoft actor-criticDirichlet policyadaptive allocation
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper frames power allocation for tracking a mobile target in an integrated sensing and communication system as a Markov decision process. It solves the problem with soft actor-critic deep reinforcement learning paired with a Dirichlet policy that outputs normalized continuous power splits under motion uncertainty. A reward function is designed to trade off tracking accuracy against communication quality and power use. Simulations indicate the resulting policy outperforms baselines on tracking while holding communication steady. The work addresses the need to share limited transmit power between sensing and data tasks when the target moves unpredictably.

Core claim

The SAC-based DRL scheme with a Dirichlet policy and carefully designed reward function enhances tracking performance compared to other baselines while sustaining communication performance in the simulated ISAC system with mobile target.

What carries the argument

Soft actor-critic reinforcement learning combined with a Dirichlet policy that generates normalized power allocation actions, steered by a reward function balancing sensing and communication metrics.

Load-bearing premise

The Markov decision process formulation and the hand-designed reward function are assumed to produce a policy that generalizes beyond the specific simulation scenarios and target-motion models used in training.

What would settle it

Deploy the trained policy on target trajectories whose motion statistics differ from the training distribution and measure whether tracking error rises markedly relative to a policy retrained on the new distribution.

Watch this falsifier. Get emailed when new claim-graph text bears on it.

Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

1 major / 1 minor

Summary. The paper models the power allocation problem in an ISAC system for tracking a mobile target as a Markov decision process and solves it using a soft actor-critic (SAC) deep reinforcement learning method combined with a Dirichlet policy for normalized continuous actions. A reward function is designed to balance tracking and communication performance. Simulation results indicate that the proposed scheme improves tracking performance compared to baselines while maintaining communication rates.

Significance. If the results are robust, the work demonstrates a practical application of DRL for adaptive resource allocation in ISAC systems, leveraging the Dirichlet policy to handle continuous normalized power allocations under random target motion. This could be valuable for systems requiring dynamic trade-offs between sensing and communication.

major comments (1)
  1. [Simulation Results] The reported improvements are obtained under fixed target-motion statistics and a hand-crafted reward function. No experiments test generalization to different process noise, turn rates, or SNR regimes, which is load-bearing for claiming that the DRL approach provides a superior policy rather than one tuned to the training distribution.
minor comments (1)
  1. The abstract mentions 'other baselines' but does not specify what they are; this should be clarified in the main text.

Simulated Author's Rebuttal

1 responses · 0 unresolved

We thank the referee for the detailed review and constructive feedback on our manuscript. Below we provide a point-by-point response to the major comment.

read point-by-point responses
  1. Referee: [Simulation Results] The reported improvements are obtained under fixed target-motion statistics and a hand-crafted reward function. No experiments test generalization to different process noise, turn rates, or SNR regimes, which is load-bearing for claiming that the DRL approach provides a superior policy rather than one tuned to the training distribution.

    Authors: We agree that the current simulations use fixed target-motion statistics and a specific reward function, and that no additional experiments vary process noise, turn rates, or SNR regimes. These parameters were selected to represent standard mobile-target ISAC tracking scenarios, and the Dirichlet policy is intended to produce normalized actions in a general manner. However, the referee is correct that broader generalization tests are needed to support claims of a superior policy beyond the training distribution. In the revised manuscript we will add simulation results under varied process noise, turn rates, and SNR conditions to address this point. revision: yes

Circularity Check

0 steps flagged

No circularity; derivation is self-contained simulation study

full rationale

The paper formulates power allocation as an MDP, applies standard SAC with a Dirichlet policy, and hand-designs a reward to trade off sensing and communication objectives. All performance claims rest on simulation comparisons to baselines under the chosen motion models and reward. No equation reduces a prediction to a fitted input by construction, no self-citation chain bears the central result, and no uniqueness theorem or ansatz is imported from prior author work. This is a normal non-circular application of DRL; the reader's score of 2 is consistent with the absence of any load-bearing circular step.

Axiom & Free-Parameter Ledger

0 free parameters · 0 axioms · 0 invented entities

Abstract-only review; no equations or implementation details are supplied, so free parameters, axioms, and invented entities cannot be enumerated. The central claim rests on an unstated assumption that the simulation environment faithfully represents real ISAC hardware and channel conditions.

reviewed 2026-06-27 · how reviews work

0 comments
Cite this review

Pith. "Pith review of Deep Reinforcement Learning for Adaptive Power Allocation in ISAC Systems with Mobile Target." pith.science (2026). https://pith.science/paper/VQPXHIJ7

@misc{pith2026260612078,
  author       = {Pith},
  title        = {Pith review of: Deep Reinforcement Learning for Adaptive Power Allocation in ISAC Systems with Mobile Target},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/VQPXHIJ7}},
  note         = {Machine review of arXiv:2606.12078}
}
Share X Bluesky LinkedIn Reddit HN
read the original abstract

In this paper, we study the power allocation for an integrated sensing and communication (ISAC) system which tracks a mobile target. We first model the problem as a Markov decision process, and then tackle it with a soft actor-critic (SAC) based deep reinforcement learning (DRL) approach. We also combine a Dirichlet policy, which naturally produces normalized continuous actions under random target motion. To exploit different features of sensing and communication operations, we carefully design a reward function such that the system can dynamically control power allocation to conserve resources. The simulation results demonstrate that the proposed scheme enhances tracking performance compared to other baselines while sustaining communication performance.

Figures

Figures reproduced from arXiv: 2606.12078 by Inkyu Lee, Jaewan Kim, Jeongwon Kim, Jihwan Moon, Sangmin Kim, Sangwon Hwang, Zhilin Fu.

Figure 1
Figure 1. Figure 1: ISAC system model A. Communication Model Denoting the path-loss factor as L ≜ [PITH_FULL_IMAGE:figures/full_fig_p002_1.png] view at source ↗
Figure 2
Figure 2. Figure 2: Feasibility of PCRB constraints across different [PITH_FULL_IMAGE:figures/full_fig_p005_2.png] view at source ↗
Figure 3
Figure 3. Figure 3: Convergence Behavior Techniques, and Applications,” Proc. IEEE., vol. 112, pp. 676–723, May 2024. [7] J. Li, X. Shao, F. Chen, S. Wan, C. Liu, Z. Wei, and D. Wing Kwan Ng, “Networked Integrated Sensing and Communications for 6G Wireless Systems,” IEEE Internet Things J., vol. 11, pp. 29062–29075, May 2024. [8] C. B. Barneto, T. Riihonen, S. D. Liyanaarachchi, M. Heino, N. Gonzalez-Prelcic, and M. Valkama, … view at source ↗

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Reference graph

Works this paper leans on

25 extracted references · 1 canonical work pages · 1 internal anchor

  1. [1]

    Enabling Joint Communication and Radar Sensing in Mobile Networks—A Survey,

    J. A. Zhang, M. L. Rahman, K. Wu, X. Huang, Y . J. Guo, S. Chen, and J. Yuan, “Enabling Joint Communication and Radar Sensing in Mobile Networks—A Survey,”IEEE Commun. Surv. Tutor ., vol. 24, pp. 306– 345, First quarter 2022

  2. [2]

    Interworking of DSRC and Cellular Network Technologies for V2X Communications: A Survey,

    K. Abboud, H. A. Omar, and W. Zhuang, “Interworking of DSRC and Cellular Network Technologies for V2X Communications: A Survey,” IEEE Trans. V eh. Technol., vol. 65, pp. 9457–9470, Dec. 2016

  3. [3]

    Framework for a Perceptive Mobile Network Using Joint Communica- tion and Radar Sensing,

    M. L. Rahman, J. A. Zhang, X. Huang, Y . J. Guo, and R. W. Heath, “Framework for a Perceptive Mobile Network Using Joint Communica- tion and Radar Sensing,”IEEE Trans. Aerosp. Electron. Syst., vol. 56, pp. 1926–1941, June 2020

  4. [4]

    IEEE 802.11ad-Based Radar: An Approach to Joint Vehicular Communication-Radar System,

    P. Kumari, J. Choi, N. Gonz ´alez-Prelcic, and R. W. Heath, “IEEE 802.11ad-Based Radar: An Approach to Joint Vehicular Communication-Radar System,”IEEE Trans. V eh. Technol., vol. 67, pp. 3012–3027, Nov. 2018

  5. [5]

    Interference Can- cellation and Iterative Detection for Orthogonal Time Frequency Space Modulation,

    P. Raviteja, K. T. Phan, Y . Hong, and E. Viterbo, “Interference Can- cellation and Iterative Detection for Orthogonal Time Frequency Space Modulation,”IEEE Trans. Wirel. Commun., vol. 17, pp. 6501–6515, Aug. 2018

  6. [6]

    Gonz ´alez-Prelcic, M

    N. Gonz ´alez-Prelcic, M. Furkan Keskin, O. Kaltiokallio, M. Valkama, D. Dardari, X. Shen, Y . Shen, M. Bayraktar, and H. Wymeersch, “The Integrated Sensing and Communication Revolution for 6G: Vision, 6 0 5000 10000 150000 10 20 30 40 50 60 70 SAC with pred. states Proposed SAC with pred. states (moving average) Proposed (moving average) (a) Total reward...

  7. [7]

    Networked Integrated Sensing and Communications for 6G Wireless Systems,

    J. Li, X. Shao, F. Chen, S. Wan, C. Liu, Z. Wei, and D. Wing Kwan Ng, “Networked Integrated Sensing and Communications for 6G Wireless Systems,”IEEE Internet Things J., vol. 11, pp. 29062–29075, May 2024

  8. [8]

    Beamformer Design and Op- timization for Joint Communication and Full-Duplex Sensing at mm- Waves,

    C. B. Barneto, T. Riihonen, S. D. Liyanaarachchi, M. Heino, N. Gonz ´alez-Prelcic, and M. Valkama, “Beamformer Design and Op- timization for Joint Communication and Full-Duplex Sensing at mm- Waves,”IEEE Trans. Commun., vol. 70, pp. 8298–8312, Dec. 2022

  9. [9]

    Integrated Sensing and Channel Estimation by Exploiting Dual Timescales for Delay-Doppler Alignment Modulation,

    Z. Xiao, Y . Zeng, F. Wen, Z. Zhang, and D. W. K. Ng, “Integrated Sensing and Channel Estimation by Exploiting Dual Timescales for Delay-Doppler Alignment Modulation,”IEEE Trans. Wirel. Commun., vol. 24, pp. 415–429, Jan. 2025

  10. [10]

    Deep Learning-Based Link Configuration for Radar-Aided Multiuser mmWave Vehicle-to-Infrastructure Communication,

    A. Graff, Y . Chen, N. Gonz ´alez-Prelcic, and T. Shimizu, “Deep Learning-Based Link Configuration for Radar-Aided Multiuser mmWave Vehicle-to-Infrastructure Communication,”IEEE Trans. V eh. Technol., vol. 72, pp. 7454–7468, Jan 2023

  11. [11]

    Throughput Maximization for UA V-Enabled Integrated Periodic Sensing and Com- munication,

    K. Meng, Q. Wu, S. Ma, W. Chen, K. Wang, and J. Li, “Throughput Maximization for UA V-Enabled Integrated Periodic Sensing and Com- munication,”IEEE Trans. Wirel. Commun., vol. 22, pp. 671–687, Jan. 2023

  12. [12]

    Asynchronous Protocol Designs for Energy Efficient Mobile Edge Computing Systems,

    S. Eom, H. Lee, J. Park, and I. Lee, “Asynchronous Protocol Designs for Energy Efficient Mobile Edge Computing Systems,”IEEE Trans. V eh. Technol., vol. 70, pp. 1013–1018, Jan. 2021

  13. [13]

    On the Design Details of SS/PBCH, Signal Generation and PRACH in 5G-NR,

    A. Chakrapani, “On the Design Details of SS/PBCH, Signal Generation and PRACH in 5G-NR,”IEEE Access, vol. 8, pp. 136617–136637, July 2020

  14. [14]

    Power allocation of integrated sensing and communication system for the internet of vehicles,

    Z. Pu, W. Wang, Z. Lao, Y . Yan, and H. Qin, “Power allocation of integrated sensing and communication system for the internet of vehicles,”IEEE Trans. Green Commun. Netw., vol. 8, pp. 1717–1728, Dec. 2024

  15. [15]

    Sensing as a Service in 6G Perceptive Networks: A Unified Framework for ISAC Resource Allocation,

    F. Dong, F. Liu, Y . Cui, W. Wang, K. Han, and Z. Wang, “Sensing as a Service in 6G Perceptive Networks: A Unified Framework for ISAC Resource Allocation,”IEEE Trans. Wirel. Commun., vol. 22, pp. 3522– 3536, Nov. 2023

  16. [16]

    EKF-Based Beamforming Design for Joint Beam Tracking and Communication Systems,

    M. Li, F. Dong, T. Liu, and F. Liu, “EKF-Based Beamforming Design for Joint Beam Tracking and Communication Systems,”IEEE Wirel. Commun. Lett., Early Access, 2025

  17. [17]

    Intelligent Resource Allocation in Joint Radar-Communication With Graph Neural Networks,

    J. Lee, Y . Cheng, D. Niyato, Y . L. Guan, and D. G. Gonz´alez, “Intelligent Resource Allocation in Joint Radar-Communication With Graph Neural Networks,”IEEE Trans. V eh. Technol., vol. 71, pp. 11120–11135, Oct. 2022

  18. [18]

    Intelli- gent Beam Tracking in Radar-Assisted MIMO-OFDM Communication Systems,

    Y . Wang, M. Lou, W. Qian, Y . Bai, L. Tang, and Y .-C. Liang, “Intelli- gent Beam Tracking in Radar-Assisted MIMO-OFDM Communication Systems,”IEEE Trans. V eh. Technol., vol. 73, pp. 16774–16789, June 2024

  19. [19]

    Cooperative Multiagent Deep Reinforcement Learning Methods for UA V-Aided Mobile Edge Computing Networks,

    M. Kim, H. Lee, S. Hwang, M. Debbah, and I. Lee, “Cooperative Multiagent Deep Reinforcement Learning Methods for UA V-Aided Mobile Edge Computing Networks,”IEEE Internet Things J., vol. 11, pp. 38040–38053, Dec. 2024

  20. [20]

    Multiagent Deep Reinforce- ment Learning for Decentralized Multi-UA V Mobile Edge Computing Networks,

    S. Hwang, H. Lee, M. Kim, and I. Lee, “Multiagent Deep Reinforce- ment Learning for Decentralized Multi-UA V Mobile Edge Computing Networks,”IEEE Internet Things J., vol. 12, pp. 14484–14497, May 2025

  21. [21]

    Radar-Assisted Predictive Beamforming for Vehicular Links: Communication Served by Sensing,

    F. Liu, W. Yuan, C. Masouros, and J. Yuan, “Radar-Assisted Predictive Beamforming for Vehicular Links: Communication Served by Sensing,” IEEE Trans Wirel. Commun., vol. 19, no. 11, pp. 7704–7719, 2020

  22. [22]

    Scaled Accuracy based Power Allocation for Multi-target Tracking with Colocated MIMO Radars,

    Y . Yuan, W. Yi, T. Kirubarajan, and L. Kong, “Scaled Accuracy based Power Allocation for Multi-target Tracking with Colocated MIMO Radars,”Signal Processing, vol. 158, pp. 227–240, Mar. 2019

  23. [23]

    Posterior Cramer-Rao Bounds for Discrete-time Nonlinear Filtering,

    P. Tichavsky, C. Muravchik, and A. Nehorai, “Posterior Cramer-Rao Bounds for Discrete-time Nonlinear Filtering,”IEEE Trans. Signal Process, vol. 46, pp. 1386–1396, May 1998

  24. [24]

    Soft Actor-Critic Algorithms and Applications

    T. Haarnoja, A. Zhou, K. Hartikainen, G. Tucker, S. Ha, J. Tan, V . Kumar, H. Zhu, A. Gupta, P. Abbeel,et al., “Soft actor-critic algorithms and applications,”arXiv preprint arXiv:1812.05905, 2018

  25. [25]

    A prescriptive Dirichlet power allocation policy with deep reinforcement learning,

    Y . Tian, M. Han, C. Kulkarni, and O. Fink, “A prescriptive Dirichlet power allocation policy with deep reinforcement learning,”Reliability Engineering & System Safety, vol. 224, p. 108529, Aug. 2022

This paper was first reviewed by grok-4.3 on June 27, 2026.