Pith. sign in

REVIEW 4 major objections 5 minor 60 references

Robust Driving Control for Autonomous Vehicles: An Intelligent General-sum Constrained Adversarial Reinforcement Learning Approach

T0 review · 4 major / 5 minor · reviewed 2026-08-04 · deepseek-v4-flash

Pith's one-line read The paper claims that training an autonomous-driving agent against a collision-seeking, multi-step adversary under two safety constraints yields a driving policy that beats state-of-the-art robust methods by at least 27.9% in success rate u

desk verdict A competent robust-driving paper with a plausible headline result, but its load-bearing C1 constraint uses the adversary's Q-function as a safety oracle that is never validated; worth refereeing, not publishing as-is. read the letter →

arxiv 2510.09041 v3 pith:V2SIDMQK submitted 2025-10-10 cs.LG cs.AI

classification cs.LGcs.AI
keywords adversarialreinforcementlearningautonomousdrivinggeneral-sumgameconstrainedpolicyoptimizationsafety-criticaleventsobservationperturbationsstrategicadversaryunprotectedleftturn
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that autonomous driving policies trained with deep reinforcement learning can be made considerably more resistant to adversarial sensor perturbations by changing both sides of adversarial training. Instead of a zero-sum adversary that minimizes the driver's reward, it proposes a strategic targeted adversary that plans multi-step perturbations and is rewarded only for causing collisions. The driver is trained under two constraints: avoid actions the adversary's own scoring deems collision-prone, and keep actions under perturbed observations close to actions under clean observations. In simulations of unprotected left turns, the resulting method keeps 100% success without attacks and beats the strongest prior baseline by 30.1 and 27.9 percentage points at two attack strengths. The intended significance is that robustness training need not sacrifice clean-environment performance or training stability.

What carries the argument

The load-bearing mechanism is the pair formed by (1) a strategic targeted adversary—a Soft Actor-Critic policy that plans multi-step attacks, emits an adversarial action, and converts it to a bounded observation perturbation via the Basic Iterative Method—and (2) a constrained driving agent solved by Lagrangian primal-dual optimization. Two constraints do the work: the collision risk constraint uses the adversary's learned Q-function as a proxy for collision probability on clean observations, and the policy consistency constraint penalizes divergence between actions on clean and perturbed observations. The general-sum reward design is what directs the adversary toward collisions instead of r

What would settle it

Run the trained agent against a stronger, unseen attacker that plans longer attack sequences or uses a larger perturbation budget than the one used in training; if the agent's success rate falls to the level of unconstrained baselines under that attacker, the claimed robustness does not generalize beyond the specific adversary used in training.

Watch

Extended reading notes

Core claim

The paper's central claim is that a general-sum adversarial game with a collision-oriented strategic adversary and a constrained learning agent, called IGCARL, produces driving policies robust to bounded observation perturbations. The adversary is a deep reinforcement learning policy that outputs adversarial actions, converted by gradient-based iterations into bounded perturbations; its reward is purely whether a collision occurs, decoupled from the driver's reward, so it targets safety-critical failures rather than efficiency loss. The agent maximizes its own reward subject to two constraints: C1 keeps the adversary's learned collision-risk value for the chosen action low on clean observati

Load-bearing premise

The method's safety guarantee rests on the assumption that the adversary's learned collision-risk scoring of actions remains accurate on clean observations even while the agent's training changes what the adversary sees; if that score is biased, the constraint will either block safe actions or allow dangerous ones.

Editorial extensions

If this is right

  • Under bounded adversarial perturbations of size 0.03 and 0.05, the trained agent achieves success rates 30.1% and 27.9% higher than the strongest prior robust method, while keeping 100% success when no attack is present.
  • The agent's actions deviate by less than 0.1 under gradient-based perturbations and less than 0.15 under random noise, indicating local policy stability beyond the specific training adversary.
  • When traffic density shifts to unseen values, the method maintains its advantage, with about 10 percentage points higher success rate than the best baseline at the largest tested perturbation.
  • The policy consistency constraint prevents the learned policy from overfitting to perturbed observations, which is why clean-environment performance does not collapse after adversarial training.
  • Training against a foresighted adversary that targets collisions reveals vulnerabilities that myopic, single-step attacks miss, suggesting robust driving policies need to be evaluated against strategic multi-step threats.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • My inference: the same two-constraint template—using an adversary-learned cost model as a safety constraint plus an action-consistency regularizer—could transfer to other safety-critical sequential decision tasks such as robot navigation or human-robot handover, though the paper only demonstrates it for driving.
  • My inference: the reported margin is measured against the specific perturbation generation procedure and scenario; how the method fares against adaptive attackers that know the constraints, or against perturbations on raw sensor inputs rather than state vectors, is untested and may be materially different.
  • My inference: the design predicts a testable trade-off—tightening the policy-consistency threshold should reduce action drift under attack but could cap performance when large perturbations push the clean action away from the optimal robust action; sweeping that threshold would reveal whether the reported operating point is on the sweet spot.
  • My inference: because the adversary is rewarded only by the collision indicator, its objective becomes identical to a safety-only agent reward; under a reward function containing only collision cost, the general-sum and zero-sum formulations would coincide, so the claimed advantage of general-sum is contingent on the agent's reward including efficiency and comfort terms.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 5 minor

Summary. The paper proposes IGCARL, an adversarial training framework for DRL-based autonomous driving in an unprotected left-turn scenario. A SAC-based strategic adversary generates targeted multi-step perturbations via BIM, while the driving agent is trained under a Lagrangian-constrained objective with two constraints: C1, which bounds the adversary's Q-value on clean actions as a proxy for collision risk, and C2, which enforces action consistency between clean and perturbed observations. Experiments in SUMO compare IGCARL with PPO, SAC, TD3, SAC-Lag, FNI, and DARRL under perturbation magnitudes ε=0.01, 0.03, 0.05, reporting success rate, collision rate, and driving efficiency. The central claim is that IGCARL outperforms DARRL by 30.1% at ε=0.03 and 27.9% at ε=0.05 in success rate.

Significance. If the empirical claim holds, the paper would make a valuable contribution to robust autonomous driving: it is the first to combine a DRL-based strategic adversary with constrained policy optimization, and it tests robustness across policy-based, gradient-based, and random perturbations, as well as generalization to different traffic densities. The writing is clear and the experimental setup is broader than many prior adversarial-RL papers. However, the headline result is weakened by two load-bearing gaps: the safety critic used in constraint C1 is the adversary's own Q-function, which is never validated as a collision-risk oracle, and the main robustness evaluation appears to use the same adversary that was co-trained with the agent, making the comparison self-referential. These issues must be resolved before the central claim can be considered fully supported.

major comments (4)
  1. [Section III.D, Eq. (14)] Constraint C1 uses E[min_i Q^adv_i(o,a)] as a collision-risk bound, but Q^adv is trained with the adversary reward r_adv = c(o,a') in Eq. (3), where a' = π(o+δ). Thus Q^adv(o,a) measures the adversary's expected return for executing its own action a, not the collision risk of the agent executing clean action a. The text's justification 'since r_adv = c(o,a)' is inconsistent with Eq. (3). No calibration or ablation is provided to show that C1 behaves as a safety critic. Please either replace C1 with an independently learned cost critic, or add an ablation without C1 to demonstrate that the reported robustness is not due to some other mechanism.
  2. [Section IV.E, Table III] The robustness evaluation appears to use the same adversary that was co-trained with the agent. If the training adversary is also the evaluation adversary, the reported success rates may reflect overfitting to that specific attacker rather than general robustness. The gradient-based and random perturbation experiments in Section IV.F are external, but they report action offsets, not success rates. Please evaluate against a held-out adversary trained with different random seeds, or against a fixed PGD attack with random restarts, and report SR under those conditions.
  3. [Section IV.E, Table III] The success-rate results are non-monotonic in the perturbation magnitude: IGCARL achieves SR 74.5 at ε=0.03 but 87.17 at ε=0.05, and DARRL similarly improves from 57.25 to 68.17. This is counterintuitive for a fixed evaluation protocol and is not explained in the text. Also, at ε=0.01 IGCARL (96.0) is lower than DARRL (97.4), which contradicts the abstract's claim of improvement 'at least 27.9%'. Please clarify whether different adversary checkpoints were used per ε, report paired confidence intervals, and qualify the claim to the conditions where it actually holds.
  4. [Section III.D, Eqs. (18)-(20)] The constrained optimization mixes expectation-level constraints with per-sample primal-dual updates. In particular, Eq. (20) updates the Lagrange multipliers using the per-sample constraint value Ck, without a convergence or feasibility analysis. In a non-stationary two-player game, this heuristic may be unstable or may not enforce the constraints at the population level. Please provide a convergence discussion, or at least report the constraint-violation curves during training to show that C1 and C2 are actually satisfied.
minor comments (5)
  1. [Abstract / Section IV.E] The phrase 'at least 27.9%' is inaccurate because at ε=0.01 IGCARL has a lower SR than DARRL. Please either report the gains for each ε separately or qualify the claim.
  2. [Section III.D, Eq. (14)] The notation in Eq. (14) is ambiguous: 'a' is not defined. It should be a = π(o), and o' should be defined as o + δ. Please state whether C1 evaluates the clean action a or the perturbed action π(o').
  3. [Table III] Several entries report standard deviation 0.00 with SR 100.00 or CR 0.00 over 200 episodes; this is possible but unusual. Please report the number of random seeds used and whether the 200 evaluation episodes are shared across methods.
  4. [Section IV.F] The gradient-based and random perturbation experiments only report action offsets, not success rates or collision rates. Adding SR under these perturbations would strengthen the claim that IGCARL is robust beyond the trained adversary.
  5. [General] Typos: 'basline' in Section IV.F, 'accleration' in Fig. 2, and 'Tmperature' in Table II. Also, reference [46] duplicates reference [41].

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: C1's Q^adv is a co-trained learned risk signal, not a fitted parameter renamed as a prediction, and the headline gains are empirical against external baselines and other perturbation types.

full rationale

I walked the paper's derivation chain. The method defines a two-player partially observable Markov game, trains a SAC-based adversary to maximize collision-return, produces perturbations by BIM gradient targeting, and trains the agent under a constrained Lagrangian objective. The C1 constraint uses the adversary's learned Q-function as a collision-risk estimator, but this is a co-trained learning signal, not an input defined in terms of the target output; the paper does not claim to derive that Q^adv is calibrated, and any calibration concern is a correctness/validation risk, not a circularity. The headline success-rate improvement is an empirical comparison against DARRL and other baselines under identical perturbation magnitudes, with additional evidence from gradient-based perturbations, random perturbations, and traffic-density generalization. The paper cites the authors' own prior work ([18], [33], [55]), but those citations are not load-bearing for the central claim: there is no uniqueness theorem, no ansatz smuggled in via self-citation, and no fitted parameter renamed as a prediction. Although the robustness evaluation uses the same adversary family used in training, this is standard adversarial-training evaluation and is supplemented by independent perturbation tests; it does not reduce the result to its inputs by construction. No equation is shown to be equivalent to another by construction, and no load-bearing premise is justified only by self-citation. Therefore the paper exhibits no significant circularity.

Assumptions & free parameters 5 free parameters · 5 assumptions · 0 invented entities

The central claim rests on standard RL convergence assumptions, a domain-specific simulation proxy, white-box access to the agent, and hand-tuned constraint thresholds. The main risk is the learned adversary Q being used as safety ground truth in C1.

free parameters (5)
  • Constraint thresholds ε1, ε2 = 0.01 each
    Hand-chosen in Eq. (14) and Table II; they determine how strongly C1 and C2 bind. No sensitivity analysis is provided.
  • Perturbation budget ε = 0.01, 0.03, 0.05
    Training/evaluation budget from Table II; the headline 27.9% gain is at ε=0.05, and the reported success rate is non-monotonic in ε (74.5 at 0.03, 87.17 at 0.05), which is not explained.
  • Dual stepsize α_λ and SAC temperature α = 5e-5, 0.1
    Chosen by tuning (Section IV.D); affects constraint satisfaction and exploration.
  • Traffic density p = 0.5 training; 0.03/0.07 generalization
    Training density is chosen by hand; generalization is tested only to two densities, not across a range.
  • BIM iterations and step size = 50 iterations, α_δ=ε/50
    Attack strength and perturbation generation details; chosen without derivation from an attack-success criterion.
assumptions (5)
  • standard math SAC and actor-critic updates converge to meaningful optima in this non-stationary two-player setting
    Equations (5)-(10) and (18)-(20) rely on standard RL convergence assumptions, but no convergence guarantee is given for the coupled adversary-agent optimization.
  • domain assumption SUMO unprotected left turn with the LC2013 model is a valid proxy for real-world driving risk
    Section IV.A; all conclusions are drawn from this single simulated scenario.
  • domain assumption The adversary has white-box access to the agent policy and its gradients for BIM
    Eqs. (12)-(13) require ∇δ||a_adv − π(o+δ)||²; this does not hold in black-box sensor deployments.
  • domain assumption Learned Q^adv is an accurate collision-risk estimator on clean observations
    Constraint C1 in Eq. (14) treats Q^adv as ground truth; no accuracy or convergence evidence is given for this non-stationary estimate.
  • ad hoc to paper Policy-consistency constraint C2 is sufficient to prevent policy drift without hurting clean performance
    C2 is taken from [27] (Section III.D); the threshold ε2=0.01 is not derived from safety requirements or an invariance principle.

how reviews work

0 comments
Cite this review

Pith. "Pith review of Robust Driving Control for Autonomous Vehicles: An Intelligent General-sum Constrained Adversarial Reinforcement Learning Approach." pith.science (2026). https://pith.science/paper/V2SIDMQK

@misc{pith2026251009041,
  author       = {Pith},
  title        = {Pith review of: Robust Driving Control for Autonomous Vehicles: An Intelligent General-sum Constrained Adversarial Reinforcement Learning Approach},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/V2SIDMQK}},
  note         = {Machine review of arXiv:2510.09041}
}
read the original abstract

Deep reinforcement learning (DRL) has demonstrated remarkable success in developing autonomous driving policies. However, its vulnerability to adversarial attacks remains a critical barrier to real-world deployment. Although existing robust methods have achieved success, they still suffer from three key issues: (i) these methods are trained against myopic adversarial attacks, limiting their abilities to respond to more strategic threats, (ii) they have trouble causing truly safety-critical events (e.g., collisions), but instead often result in minor consequences, and (iii) these methods can introduce learning instability and policy drift during training due to the lack of robust constraints. To address these issues, we propose Intelligent General-sum Constrained Adversarial Reinforcement Learning (IGCARL), a novel robust autonomous driving approach that consists of a strategic targeted adversary and a robust driving agent. The strategic targeted adversary is designed to leverage the temporal decision-making capabilities of DRL to execute strategically coordinated multi-step attacks. In addition, it explicitly focuses on inducing safety-critical events by adopting a general-sum objective. The robust driving agent learns by interacting with the adversary to develop a robust autonomous driving policy against adversarial attacks. To ensure stable learning in adversarial environments and to mitigate policy drift caused by attacks, the agent is optimized under a constrained formulation. Extensive experiments show that IGCARL improves the success rate by at least 27.9% over state-of-the-art methods, demonstrating superior robustness to adversarial attacks and enhancing the safety and reliability of DRL-based autonomous driving.

Figures

Figures reproduced from arXiv: 2510.09041 by the authors.

Figure 1
Figure 1. Overview of IGCARL and its role in addressing key challenges. IGCARL consists of two main components: a strategic targeted adversary designed to find critical vulnerabilities, and a robust driving agent trained to withstand the adversary and produce a robust driving policy. B. Robust DRL-based AD Methods Against Adversarial At￾tacks The deployment of DRL in AD has raised critical safety concerns. Due to the black-bo… view at source ↗
Figure 2
Figure 2. Qualitative analysis of constraints (C1 and C2) in an unprotected left turn scenario. The top row shows C1: Without C1 (left), the agent selects a hazardous acceleration leading to a collision. With C1 (right), the agent brakes safely. The bottom row shows C2: without C2 (left), the safe braking is overridden by a final acceleration; with C2 (right), the braking is preserved, preventing a collision. where αδ is the … view at source ↗
Figure 3
Figure 3. Visualization of action distributions under clean and adversarial environments. Subfigures (a)–(e) correspond to SAC, SAC Lag, FNI, DARRL, and IGCARL, respectively. Column (1) shows the action distribution under clean observations, while Column (2) shows the shift in the action distribution caused by adversarial perturbations, denoted as ∆ = π(o ′ ) − π(o) [PITH_FULL_IMAGE:figures/full_fig_p009_3.png] view at source ↗
Figures from the paper (3 more)
Figure 4
Figure 4. Figure 4: Sensitivity analysis of policies under gradient-based perturbations. The y-axis denotes the perturbation along the action gradient direction, while the x-axis represents the perturbation along the direction orthogonal to the action gradient. that minor perturbations ca…
Figure 5
Figure 5. Figure 5: Action offset under random noise perturbations with ϵ = 0.05. 3) Random perturbations: We further evaluate the robust￾ness of the methods by introducing random noise. This de￾sign is motivated by several considerations: (i) random noise can simulate small, unpredictabl…
Figure 6
Figure 6. Figure 6: Performance comparison under different traffic densities. (a)–(c) show results with ϵ = 0.01, 0.03, and 0.05 in Env1 (p = 0.03), while (d)–(f) present the corresponding results in Env2 (p = 0.07). REFERENCES [1] J. Zhao, Y. Wu, R. Deng, S. Xu, J. Gao, and A. Burke, “A …

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

60 extracted references · 2 linked inside Pith

  1. [1]

    A survey of autonomous driving from a deep learning perspective,

    J. Zhao, Y . Wu, R. Deng, S. Xu, J. Gao, and A. Burke, “A survey of autonomous driving from a deep learning perspective,”ACM Comput. Surv., vol. 57, no. 10, pp. 1–60, 2025

  2. [2]

    Sparsedrive: End-to-end autonomous driving via sparse scene representation,

    W. Sun, X. Lin, Y . Shi, C. Zhang, H. Wu, and S. Zheng, “Sparsedrive: End-to-end autonomous driving via sparse scene representation,” in2025 IEEE International Conference on Robotics and Automation (ICRA). IEEE, 2025, pp. 8795–8801

  3. [3]

    Embodied ai-enhanced vehicular networks: An integrated vision language models and reinforcement learning method,

    R. Zhang, C. Zhao, H. Du, D. Niyato, J. Wang, S. Sawadsitang, X. Shen, and D. I. Kim, “Embodied ai-enhanced vehicular networks: An integrated vision language models and reinforcement learning method,” IEEE Trans. on Mob. Comput., 2025

  4. [4]

    Deepseek-r1 incentivizes reasoning in llms through reinforcement learning,

    D. Guo, D. Yang, H. Zhang, J. Song, P. Wang, Q. Zhu, R. Xu, R. Zhang, S. Ma, X. Biet al., “Deepseek-r1 incentivizes reasoning in llms through reinforcement learning,”Nature, vol. 645, no. 8081, pp. 633–638, 2025

  5. [5]

    A Survey on Recent Advancements in Autonomous Driving Using Deep Reinforce- ment Learning: Applications, Challenges, and Solutions,

    R. Zhao, Y . Li, Y . Fan, F. Gao, M. Tsukada, and Z. Gao, “A Survey on Recent Advancements in Autonomous Driving Using Deep Reinforce- ment Learning: Applications, Challenges, and Solutions,”IEEE Trans. on Intell. Transp. Syst., vol. 25, no. 12, pp. 19 365–19 398, Dec. 2024

  6. [6]

    Self-Learned Autonomous Driving at Unsignalized Intersections: A Hierarchical Reinforced Learning Approach for Feasible Decision-Making,

    M. Al-Sharman, R. Dempster, M. A. Daoud, M. Nasr, D. Rayside, and W. Melek, “Self-Learned Autonomous Driving at Unsignalized Intersections: A Hierarchical Reinforced Learning Approach for Feasible Decision-Making,”IEEE Trans. on Intell. Transp. Syst., vol. 24, no. 11, pp. 12 345–12 356, Nov. 2023

  7. [7]

    Integration of planning and deep reinforcement learning in speed and lane change decision-making for highway autonomous driving,

    S. Zhang, W. Zhuang, B. Li, K. Li, T. Xia, and B. Hu, “Integration of planning and deep reinforcement learning in speed and lane change decision-making for highway autonomous driving,”IEEE Trans. on Transp. Electrification, vol. 11, no. 1, pp. 521–535, 2024

  8. [8]

    Highway autonomous vehicle decision-making method based on prior knowledge and improved ex- perience replay reinforcement learning algorithm,

    Z. Wang, P. Li, Z. Wang, and Z. Li, “Highway autonomous vehicle decision-making method based on prior knowledge and improved ex- perience replay reinforcement learning algorithm,”Expert Syst. with Appl., p. 127927, 2025

Show all 60 references
  1. [9]

    Reinforcement Learning- Based Multi-Lane Cooperative Control for On-Ramp Merging in Mixed- Autonomy Traffic,

    L. Liu, X. Li, Y . Li, J. Li, and Z. Liu, “Reinforcement Learning- Based Multi-Lane Cooperative Control for On-Ramp Merging in Mixed- Autonomy Traffic,”IEEE Internet Things J., pp. 1–1, 2024

  2. [10]

    On-Ramp Merging for Highway Autonomous Driving: An Application of a New Safety Indicator in Deep Reinforcement Learning,

    G. Li, W. Zhou, S. Lin, S. Li, and X. Qu, “On-Ramp Merging for Highway Autonomous Driving: An Application of a New Safety Indicator in Deep Reinforcement Learning,”Automot. Innov., vol. 6, no. 3, pp. 453–465, Aug. 2023

  3. [11]

    Ensemble Quantile Networks: Uncertainty-Aware Reinforcement Learning with Applications in Au- tonomous Driving,

    C.-J. Hoel, K. Wolff, and L. Laine, “Ensemble Quantile Networks: Uncertainty-Aware Reinforcement Learning with Applications in Au- tonomous Driving,”IEEE Trans. on Intell. Transp. Syst., vol. 24, no. 6, pp. 6030–6041, Jun. 2023

  4. [12]

    Predictive trajectory planning for autonomous vehicles at intersections using reinforcement learning,

    E. Zhang, R. Zhang, and N. Masoud, “Predictive trajectory planning for autonomous vehicles at intersections using reinforcement learning,” Transp. Research Part C: Emerg. Technol., vol. 149, p. 104063, Apr. 2023

  5. [13]

    Human as ai mentor: En- hanced human-in-the-loop reinforcement learning for safe and efficient autonomous driving,

    Z. Huang, Z. Sheng, C. Ma, and S. Chen, “Human as ai mentor: En- hanced human-in-the-loop reinforcement learning for safe and efficient autonomous driving,”Commun. Transp. Research, vol. 4, p. 100127, 2024

  6. [14]

    Decision-making of autonomous vehicles in interactions with jaywalk- ers: A risk-aware deep reinforcement learning approach,

    Z. Zhang, H. Li, T. Chen, N. Sze, W. Yang, Y . Zhang, and G. Ren, “Decision-making of autonomous vehicles in interactions with jaywalk- ers: A risk-aware deep reinforcement learning approach,”Accident Anal. & Prev., vol. 210, p. 107843, Feb. 2025

  7. [15]

    Uncertainty quantification for safe and reliable autonomous vehicles: A review of methods and applications,

    K. Wang, C. Shen, X. Li, and J. Lu, “Uncertainty quantification for safe and reliable autonomous vehicles: A review of methods and applications,”IEEE Trans. on Intell. Transp. Syst., 2025

  8. [16]

    Seeing is not believing: Robust reinforcement learning against spurious correlation,

    W. Ding, L. Shi, Y . Chi, and D. Zhao, “Seeing is not believing: Robust reinforcement learning against spurious correlation,”Adv. Neural Inf. Process. Syst., vol. 36, pp. 66 328–66 363, 2023

  9. [17]

    Adversarial Machine Learning Attacks and Defences in Multi-Agent Reinforcement Learning,

    M. Standen, J. Kim, and C. Szabo, “Adversarial Machine Learning Attacks and Defences in Multi-Agent Reinforcement Learning,”ACM Comput. Surv., vol. 57, no. 5, pp. 1–35, May 2025

  10. [18]

    Less is more: A stealthy and efficient adversarial attack method for drl-based autonomous driving policies,

    J. Fan, X. Lei, X. Chang, J. Mi ˇsi´c, and V . B. Mi ˇsi´c, “Less is more: A stealthy and efficient adversarial attack method for drl-based autonomous driving policies,”arXiv preprint arXiv:2412.03051, 2024

  11. [19]

    Enhancing cyber-resilience in integrated energy system scheduling with demand response using deep reinforcement learning,

    Y . Li, W. Ma, Y . Li, S. Li, Z. Chen, and M. Shahidehpour, “Enhancing cyber-resilience in integrated energy system scheduling with demand response using deep reinforcement learning,”Appl. Energy, vol. 379, p. 124831, 2025

  12. [20]

    Adversarial training with anti- adversaries,

    X. Zhou, O. Wu, and N. Yang, “Adversarial training with anti- adversaries,”IEEE Trans. on Pattern Anal. Mach. Intell., 2024

  13. [21]

    Robust Decision Making for Au- tonomous Vehicles at Highway On-Ramps: A Constrained Adversarial Reinforcement Learning Approach,

    X. He, B. Lou, H. Yang, and C. Lv, “Robust Decision Making for Au- tonomous Vehicles at Highway On-Ramps: A Constrained Adversarial Reinforcement Learning Approach,”IEEE Trans. on Intell. Transp. Syst., vol. 24, no. 4, pp. 4103–4113, Apr. 2023

  14. [22]

    Robust lane change decision for autonomous vehicles in mixed traffic: A safety-aware multi- agent adversarial reinforcement learning approach,

    T. Wang, M. Ma, S. Liang, J. Yang, and Y . Wang, “Robust lane change decision for autonomous vehicles in mixed traffic: A safety-aware multi- agent adversarial reinforcement learning approach,”Transp. Research Part C: Emerg. Technol., vol. 172, p. 105005, Mar. 2025

  15. [23]

    Robust Training in Multiagent Deep Reinforcement Learning Against Optimal JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 12 Adversary,

    W. Guo, G. Liu, Z. Zhou, J. Wang, Y . Tang, and M. Wang, “Robust Training in Multiagent Deep Reinforcement Learning Against Optimal JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021 12 Adversary,”IEEE Trans. on Syst. Man, Cybern. Syst., vol. 55, no. 7, pp. 4957–4968, Jul. 2025

  16. [24]

    Rat: Adversarial attacks on deep reinforcement agents for targeted behaviors,

    F. Bai, R. Liu, Y . Du, Y . Wen, and Y . Yang, “Rat: Adversarial attacks on deep reinforcement agents for targeted behaviors,” inProceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 15, 2025, pp. 15 453–15 461

  17. [25]

    Robust Lane Change Decision Making for Autonomous Vehicles: An Observation Adversarial Rein- forcement Learning Approach,

    X. He, H. Yang, Z. Hu, and C. Lv, “Robust Lane Change Decision Making for Autonomous Vehicles: An Observation Adversarial Rein- forcement Learning Approach,”IEEE Trans. on Intell. Veh., vol. 8, no. 1, pp. 184–193, Jan. 2023

  18. [26]

    Explainable Deep Adversarial Reinforcement Learning Approach for Robust Autonomous Driving,

    C. Wang and N. Aouf, “Explainable Deep Adversarial Reinforcement Learning Approach for Robust Autonomous Driving,”IEEE Trans. on Intell. Veh., pp. 1–13, 2024

  19. [27]

    Trustworthy autonomous driving via defense-aware robust reinforcement learning against worst-case obser- vational perturbations,

    X. He, W. Huang, and C. Lv, “Trustworthy autonomous driving via defense-aware robust reinforcement learning against worst-case obser- vational perturbations,”Transp. Research Part C: Emerg. Technol., vol. 163, p. 104632, Jun. 2024

  20. [28]

    Improved ro- bustness and safety for autonomous vehicle control with adversarial reinforcement learning,

    X. Ma, K. Driggs-Campbell, and M. J. Kochenderfer, “Improved ro- bustness and safety for autonomous vehicle control with adversarial reinforcement learning,” in2018 IEEE Intelligent Vehicles Symposium (IV). IEEE, 2018, pp. 1665–1671

  21. [29]

    CAT: Closed-loop Adversarial Training for Safe End-to-End Driving,

    L. Zhang, Z. Peng, Q. Li, and B. Zhou, “CAT: Closed-loop Adversarial Training for Safe End-to-End Driving,” Oct. 2023

  22. [30]

    Adversarial Stress Test for Autonomous Vehicle Via Series Reinforcement Learning Tasks With Reward Shaping,

    X. Cai, X. Bai, Z. Cui, P. Hang, H. Yu, and Y . Ren, “Adversarial Stress Test for Autonomous Vehicle Via Series Reinforcement Learning Tasks With Reward Shaping,”IEEE Trans. on Intell. Veh., pp. 1–16, 2024

  23. [31]

    Targeted Attack on Deep RL-based Autonomous Driving with Learned Visual Patterns,

    P. Buddareddygari, T. Zhang, Y . Yang, and Y . Ren, “Targeted Attack on Deep RL-based Autonomous Driving with Learned Visual Patterns,” in2022 International Conference on Robotics and Automation (ICRA), May 2022, pp. 10 571–10 577

  24. [32]

    Dynamic Adversarial Attacks on Autonomous Driving Systems,

    A. Chahe, C. Wang, A. Jeyapratap, K. Xu, and L. Zhou, “Dynamic Adversarial Attacks on Autonomous Driving Systems,” inRobotics: Science and Systems XX. Robotics: Science and Systems Foundation, Jul. 2024

  25. [33]

    Sharpening the spear: Adaptive expert- guided adversarial attack against drl-based autonomous driving policies,

    J. Fan, X. Lei, and X. Chang, “Sharpening the spear: Adaptive expert- guided adversarial attack against drl-based autonomous driving policies,” arXiv preprint arXiv:2506.18304, 2025

  26. [34]

    Stealthy and Efficient Adversarial Attacks against Deep Reinforcement Learning,

    J. Sun, T. Zhang, X. Xie, L. Ma, Y . Zheng, K. Chen, and Y . Liu, “Stealthy and Efficient Adversarial Attacks against Deep Reinforcement Learning,”Proc. AAAI Conf. on Artif. Intell., vol. 34, no. 04, pp. 5883–5891, Apr. 2020

  27. [35]

    Refining the black-box AI optimization with CMA-ES and ORM in the energy management for fuel cell electric vehicles,

    J. Hu, J. Li, M. Liu, Y . Huang, Q. Zhou, Y . Liu, Z. Chen, J. Yang, J. Jiang, and Y . Zhang, “Refining the black-box AI optimization with CMA-ES and ORM in the energy management for fuel cell electric vehicles,”Energy Convers. Manag., vol. 325, p. 119399, Feb. 2025

  28. [36]

    Toward Trustworthy Decision-Making for Autonomous Vehicles: A Robust Reinforcement Learning Approach with Safety Guarantees,

    X. He, W. Huang, and C. Lv, “Toward Trustworthy Decision-Making for Autonomous Vehicles: A Robust Reinforcement Learning Approach with Safety Guarantees,”Engineering, vol. 33, pp. 77–89, Feb. 2024

  29. [37]

    Attention-Based Highway Safety Planner for Autonomous Driving via Deep Reinforcement Learning,

    G. Chen, Y . Zhang, and X. Li, “Attention-Based Highway Safety Planner for Autonomous Driving via Deep Reinforcement Learning,”IEEE Trans. on Veh. Technol., vol. 73, no. 1, pp. 162–175, Jan. 2024

  30. [38]

    Autonomous Driving using Safe Reinforcement Learning by Incorporating a Regret-based Human Lane-Changing Decision Model,

    D. Chen, L. Jiang, Y . Wang, and Z. Li, “Autonomous Driving using Safe Reinforcement Learning by Incorporating a Regret-based Human Lane-Changing Decision Model,” in2020 American Control Conference (ACC), Jul. 2020, pp. 4355–4361

  31. [39]

    Trustworthy Human-AI Collab- oration: Reinforcement Learning with Human Feedback and Physics Knowledge for Safe Autonomous Driving,

    Z. Huang, Z. Sheng, and S. Chen, “Trustworthy Human-AI Collab- oration: Reinforcement Learning with Human Feedback and Physics Knowledge for Safe Autonomous Driving,” Sep. 2024

  32. [40]

    Human- Guided Deep Reinforcement Learning for Optimal Decision Making of Autonomous Vehicles,

    J. Wu, H. Yang, L. Yang, Y . Huang, X. He, and C. Lv, “Human- Guided Deep Reinforcement Learning for Optimal Decision Making of Autonomous Vehicles,”IEEE Trans. on Syst. Man, Cybern. Syst., vol. 54, no. 11, pp. 6595–6609, Nov. 2024

  33. [42]

    Large Language Model guided Deep Reinforcement Learning for Decision Making in Autonomous Driving,

    H. Pang, Z. Wang, and G. Li, “Large Language Model guided Deep Reinforcement Learning for Decision Making in Autonomous Driving,” Dec. 2024

  34. [43]

    Efficient Learning of Safe Driving Policy via Human-AI Copilot Optimization,

    Q. Li, Z. Peng, and B. Zhou, “Efficient Learning of Safe Driving Policy via Human-AI Copilot Optimization,” Feb. 2022

  35. [44]

    Optimizing Autonomous Driving for Safety: A Human-Centric Approach with LLM-Enhanced RLHF,

    Y . Sun, N. Salami Pargoo, P. Jin, and J. Ortiz, “Optimizing Autonomous Driving for Safety: A Human-Centric Approach with LLM-Enhanced RLHF,” inCompanion of the 2024 on ACM International Joint Con- ference on Pervasive and Ubiquitous Computing. Melbourne VIC Australia: ACM, Oc...

  36. [45]

    Robust Adversarial Reinforcement Learning,

    L. Pinto, J. Davidson, R. Sukthankar, and A. Gupta, “Robust Adversarial Reinforcement Learning,” inProceedings of the 34th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, D. Precup and Y . W. Teh, Eds., vol. 70. PMLR, Aug. 2017, pp....

  37. [46]

    Safety-Aware Human-in-the- Loop Reinforcement Learning With Shared Control for Autonomous Driving,

    W. Huang, H. Liu, Z. Huang, and C. Lv, “Safety-Aware Human-in-the- Loop Reinforcement Learning With Shared Control for Autonomous Driving,”IEEE Trans. on Intell. Transp. Syst., vol. 25, no. 11, pp. 16 181–16 192, Nov. 2024

  38. [47]

    Adversarial examples in the physical world,

    A. Kurakin, I. J. Goodfellow, and S. Bengio, “Adversarial examples in the physical world,” inArtificial intelligence safety and security. Chapman and Hall/CRC, 2018, pp. 99–112

  39. [48]

    Mi- croscopic Traffic Simulation using SUMO,

    P. A. Lopez, M. Behrisch, L. Bieker-Walz, J. Erdmann, Y .-P. Fl ¨otter¨od, R. Hilbrich, L. L ¨ucken, J. Rummel, P. Wagner, and E. Wiessner, “Mi- croscopic Traffic Simulation using SUMO,” in2018 21st International Conference on Intelligent Transportation Systems (ITSC), Nov. 20...

  40. [49]

    Unprotected Left- Turn Behavior Model Capturing Path Variations at Intersections,

    J. Zhao, V . L. Knoop, J. Sun, Z. Ma, and M. Wang, “Unprotected Left- Turn Behavior Model Capturing Path Variations at Intersections,”IEEE Trans. on Intell. Transp. Syst., vol. 24, no. 9, pp. 9016–9030, Sep. 2023

  41. [50]

    Self-learned autonomous driving at unsignalized intersec- tions: A hierarchical reinforced learning approach for feasible decision- making,

    M. Al-Sharman, R. Dempster, M. A. Daoud, M. Nasr, D. Rayside, and W. Melek, “Self-learned autonomous driving at unsignalized intersec- tions: A hierarchical reinforced learning approach for feasible decision- making,”IEEE Trans. on Intell. Transp. Syst., vol. 24, no. 11, pp. 1...

  42. [51]

    Sumo’s lane-changing model,

    J. Erdmann, “Sumo’s lane-changing model,” inModeling Mobility with Open Data: 2nd SUMO Conference 2014 Berlin, Germany, May 15-16,

  43. [52]

    Proximal Policy Optimization Algorithms,

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal Policy Optimization Algorithms,” Aug. 2017

  44. [53]

    Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor,

    T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine, “Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor,” inProceedings of the 35th International Conference on Machine Learning. PMLR, Jul. 2018, pp. 1861–1870

  45. [54]

    Addressing Function Approx- imation Error in Actor-Critic Methods,

    S. Fujimoto, H. Hoof, and D. Meger, “Addressing Function Approx- imation Error in Actor-Critic Methods,” inProceedings of the 35th International Conference on Machine Learning. PMLR, Jul. 2018, pp. 1587–1596

  46. [55]

    Energy- Constrained Safe Path Planning for UA V-Assisted Data Collection of Mobile IoT Devices,

    J. Fan, X. Chang, J. Mi ˇsi´c, V . B. Miˇsi´c, T. Yang, and Y . Gong, “Energy- Constrained Safe Path Planning for UA V-Assisted Data Collection of Mobile IoT Devices,”IEEE Internet Things J., pp. 1–1, 2024

  47. [56]

    T-td3: A reinforcement learning framework for stable grasping of deformable objects using tactile prior,

    Y . Zhou, Y . Jin, P. Lu, S. Jiang, Z. Wang, and B. He, “T-td3: A reinforcement learning framework for stable grasping of deformable objects using tactile prior,”IEEE Trans. on Autom. Sci. Eng., 2024

  48. [57]

    Visionary Policy Iteration for Continuous Control,

    B. Dong, L. Huang, X. Ma, H. Chen, and W. Zhang, “Visionary Policy Iteration for Continuous Control,”IEEE Trans. on Syst. Man, Cybern. Syst., vol. 55, no. 4, pp. 2707–2720, Apr. 2025

  49. [58]

    Resource allocation for uav-enabled spatio-temporal crowdsourcing in smart cities,

    Y . Liu, W. Mao, X. Li, W. Huangfu, Y . Xiao, Y . Ji, H. Zhang, and K. Long, “Resource allocation for uav-enabled spatio-temporal crowdsourcing in smart cities,”IEEE Trans. on Veh. Technol., 2025

  50. [59]

    Learning to walk in the real world with minimal human effort,

    S. Ha, P. Xu, Z. Tan, S. Levine, and J. Tan, “Learning to walk in the real world with minimal human effort,” inConference on Robot Learning. PMLR, 2021, pp. 1110–1120

  51. [60]

    Fear-Neuro-Inspired Reinforcement Learning for Safe Autonomous Driving,

    X. He, J. Wu, Z. Huang, Z. Hu, J. Wang, A. Sangiovanni-Vincentelli, and C. Lv, “Fear-Neuro-Inspired Reinforcement Learning for Safe Autonomous Driving,”IEEE Trans. on Pattern Anal. Mach. Intell., vol. 46, no. 1, pp. 267–279, Jan. 2024

  52. [2014]

    Springer, 2015, pp. 105–123

Pith tools

Reviewed August 4, 2026 · model on record in the stance chip above.