REVIEW 4 major objections 5 minor 2 cited by
Action Space Reduction Strategies for Reinforcement Learning in Autonomous Driving
T0 review · 4 major / 5 minor · reviewed 2026-08-06 · deepseek-v4-flash
Pith's one-line read A PPO-based driving agent converges about twice as fast when its steering commands are limited to a state-dependent window of five values rather than the full discrete range.
desk verdict The paper's headline speedup for relative action reduction is confounded with a change in action representation, and the dynamic masking variant is slower than baseline in the paper's own table, so the central claim isn't supported, though the comparison is worth refereeing. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The machinery is a state-conditioned binary mask over a flattened discrete steering-throttle action set. For an agent with current steering index $s_t$, the valid indices are $\{i : |i-s_t| \le 2, i \in I_f\}$, so at most five steering values remain available at each step, while throttle stays in $\{0, 0.2\}$. The mask is applied through a maskable implementation of PPO, keeping the full action dimension constant so that the index-to-action mapping stays stable. In the relative variant, the steering command is re-parameterised as an adjustment $s_t = s_{t-1} + \Delta s$ with $\Delta s \in \{-0.2, -0.1, 0, 0.1, 0.2\}$, and boundary masks prevent adjustments that would push the steering angle outside its valid range. Together these two devices cut the branching factor from 11 or 21 to 5 while preserving the input/output interface of the policy.
What would settle it
Repeat the evaluation with, say, 50 seeded episodes per route and construct confidence intervals for success rate; if the relative ±0.5 variant's advantage over the full ±0.5 baseline shrinks to overlap, the paper's efficiency claim is not supported. As a second check, retrain each configuration with multiple seeds and see whether the ordering of convergence steps (Rel-0.5 fastest, then Fix-21012, Rel-1.0, F-0.5, Dyn-0.5 slowest) reproduces.
Extended reading notes
Core claim
The central claim is that state-conditioned action-space pruning changes how a PPO driving agent learns, not just which actions it is allowed to take. With five steering options ($\pm 0.2$ around the current steering index) and two throttle values, the agent explores far fewer irrelevant commands while still being able to make fine adjustments. The paper finds that the relative variant, which re-parameterises steering as a bounded adjustment $s_t = s_{t-1} + \Delta s$, converges around two times faster than the full $\pm 0.5$ baseline, achieves 100% success on straight routes versus 50% for the baseline, and posts the highest combined efficiency score on one-turn and full-route tests, despite lower raw reward than the full baseline in some cases. Dynamic masking keeps the full action set and masks invalid choices, which gives the most stable lane-keeping behavior but a slower step-to-target. These results are attributed to the combination of a smaller branching factor and a stable mapping between action indices and their effects, as measured under the paper's evaluation protocol of five deterministic episodes per test route.
Load-bearing premise
Every success-rate conclusion rests on five deterministic episodes per test route, so a difference of 25% versus 75% on one-turn routes could easily flip with a different set of seeds.
Editorial extensions
If this is right
- Relative ±0.5 reaches the 60%-of-max reward threshold in about 1.02 million steps, roughly 1.84 times faster than full ±0.5, so training cost drops by roughly half for comparable route-completion rates.
- On straight routes, relative ±0.5 and dynamic ±0.5 complete 100% of five deterministic runs, versus 50% for full ±0.5, suggesting the pruned spaces are not simply trading away difficulty.
- Wider-range variants (full ±1, dynamic ±1, relative ±1) converge slower, reach lower final reward, and show higher lane deviation, so the benefit appears tied to keeping the action window tight.
- There is no single best configuration across all metrics: full ±0.5 has the highest reward on the full route, while relative ±0.5 has the best efficiency score, which implies action-space design should be tuned for the route complexity.
Reading between the lines
- A direct extension would push the same two mechanisms into simpler continuous-control benchmarks: if bounded relative deltas speed up learning there too, the effect is due to the re-parameterisation rather than to driving-specific kinematic constraints.
- One testable stress test is to increase the evaluation budget from five to many seeded episodes per route; if the gap between Dyn-0.5's 25% and Rel-0.5's 70% success on one-turn routes narrows, part of the claimed ranking is small-sample noise.
- The paper leaves open whether a learned predictor of valid steering windows could replace the hand-designed $\pm 0.2$ mask; comparing a trained mask against the fixed window would separate the benefit of context-awareness from the benefit of a tight window.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes two state-conditioned action-space modification strategies for discrete-action PPO in the CARLA driving simulator: dynamic masking, which restricts steering to a window around the current steering index, and relative reduction, which re-parameterizes steering as bounded increments from the current angle. It compares nine configurations (full, fixed, dynamic, relative) over 4M training steps and reports training curves, convergence steps, and test results on five selected configurations across route types. The main claimed finding is that action-space reduction, especially the relative +/-0.5 variant, improves convergence speed and maintains success rates relative to a full action space.
Significance. The paper targets a real problem: sample inefficiency caused by large discrete action spaces in RL-based driving. The systematic comparison of nine action-space configurations, with a clearly specified masking procedure and an explicit convergence metric, is a useful empirical contribution. The relative-action representation for steering is a plausible way to reduce the branching factor while preserving fine control. However, the evidence as presented does not yet establish the central claim: the dynamic masking method converges slower than the full baseline, the relative method's gains are confounded with a change in MDP representation, and all quantitative results rest on single training runs and very small evaluation samples. If the authors add multi-seed results, proper control conditions, and error bars, the comparison would be a valuable reference for practitioners.
major comments (4)
- [Section V-A and Table V] Dyn-0.5, one of the two proposed strategies, reaches the convergence threshold at 2,752,512 steps, which is 1.46x slower than the F-0.5 full baseline (1,884,160 steps). This contradicts the abstract's claim that the proposed dynamic and relative schemes achieve a favorable balance between learning speed, control precision, and generalization. The authors should either narrow the claim to the relative scheme or provide a mechanistic explanation for why dynamic masking slows convergence despite reducing the effective branching factor from 22 to at most 10 actions.
- [Section III-A2, Eq. (2), and Table V] Rel-0.5's faster convergence cannot be attributed to action-space reduction alone because Eq. (2) changes the action semantics from absolute steering targets to relative increments, which changes the transition dynamics of the MDP. The comparison with F-0.5 therefore conflates reduction with representation change. A control condition using the same five absolute steering values without the relative re-parameterization, or a relative variant with the full delta set, is needed to separate the two effects.
- [Section IV-B and Tables III-V] The evaluation uses a single training run per configuration and only five deterministic episodes per test route, with no error bars, confidence intervals, or significance tests. The success-rate differences in Table III (e.g., 25% vs 75% on one-turn routes) are particularly vulnerable to noise, and the decision to evaluate only the five configurations that looked best in training (Section V-A) introduces selection bias. The authors should report multiple seeds, per-route variance, and ideally a pre-registered evaluation of all nine configurations.
- [Section IV-B2] The convergence threshold Rtgt = 0.6 Rmax is a relative threshold, and Rmax is the maximum reward achieved by any method in a single training run. This makes the step-to-target and the efficiency metric E sensitive to one noisy high-reward run and to the arbitrary choice of p = 0.6. The authors should report sensitivity of the ranking to p and to Rmax estimation, and report the reward curves of the four configurations excluded from Table V instead of assigning them a convergence rate of zero by fiat.
minor comments (5)
- [Section IV-B1] The word 'combinaion' should be 'combination' in the list of evaluation metrics.
- [Section III-A1 and III-A2] The symbol st is used both for the steering index in Eq. (1) and for the steering angle in Eq. (2); using distinct symbols would avoid confusion.
- [Section IV-B2] The efficiency formula appears to use 10^-7, but the numerical values in Table III are consistent with 10^7; the exponent should be corrected.
- [Section V-B] In the One-Turn Scenarios paragraph, 'The Full has the highest success rate' should say 'F-0.5' to match the configuration labels in Table II.
- [Fig. 1 caption] The caption says 'Routes B are the 4 different parts' but the figure labels the training route as A; the wording should clarify that route A is divided into four sections labeled B, and the test routes are labeled C through L.
Circularity Check
No significant circularity: the paper is an empirical benchmark study whose claims rest on direct training and test measurements, not on self-referential derivations or self-citations.
full rationale
We walked the paper's claimed derivation chain. The central claims—that dynamic masking and relative reduction improve training efficiency or strike a favorable balance—are empirical findings from CARLA experiments, not predictions derived from a model. The convergence metric in Section IV-B2 defines R_tgt = 0.6 * Rmax, where Rmax is the highest observed reward among the compared methods; this is a relative benchmark normalization, not an equation whose output is assumed by construction. No parameter is fitted to a subset of data and then reported as a prediction on a closely related quantity. The action-space configurations are defined independently in Sections III-A and III-B, and Table V reports step-to-target values read off the training curves rather than quantities algebraically forced by the definitions. The paper cites no prior work by the same authors as load-bearing support, invokes no uniqueness theorem, and does not smuggle an ansatz in via citation. A possible confound is that Rel-0.5 re-parameterizes steering as bounded deltas, so its speed-up may not be solely due to reduction; that is a threat to attribution and validity, not circularity. Likewise, the small number of test episodes affects statistical reliability, but it does not make the argument circular. Under the hard rules, no circular step can be exhibited with a quotation showing a specific reduction of a result to its own inputs, so the appropriate finding is no significant circularity.
Assumptions & free parameters
free parameters (5)
- dynamic mask window half-width =
2 steering indices
- relative steering increment step =
0.1
- relative steering adjustment bound =
+/-0.2
- throttle action set =
{0, 0.2}
- convergence threshold fraction p =
0.6
assumptions (4)
- domain assumption CARLA Town07 with synchronous mode faithfully represents real-world driving conditions for evaluating autonomous driving policies.
- domain assumption The reward function in Section IV-A2 correctly encodes desirable driving behavior (lane keeping, speed, progress, collision avoidance).
- domain assumption The MobileNetV3 encoder pretrained on natural images provides a suitable feature representation for BEV semantic frames without task-specific fine-tuning.
- domain assumption The four training route sections and ten novel test routes in Town07 are representative of general urban driving.
Cite this review
Pith. "Pith review of Action Space Reduction Strategies for Reinforcement Learning in Autonomous Driving." pith.science (2026). https://pith.science/paper/FRXIF2VT
@misc{pith2026250705251,
author = {Pith},
title = {Pith review of: Action Space Reduction Strategies for Reinforcement Learning in Autonomous Driving},
year = {2026},
howpublished = {\url{https://pith.science/paper/FRXIF2VT}},
note = {Machine review of arXiv:2507.05251}
}
read the original abstract
Reinforcement Learning (RL) offers a promising framework for autonomous driving by enabling agents to learn control policies through interaction with environments. However, large and high-dimensional action spaces often used to support fine-grained control can impede training efficiency and increase exploration costs. In this study, we introduce and evaluate two novel structured action space modification strategies for RL in autonomous driving: dynamic masking and relative action space reduction. These approaches are systematically compared against fixed reduction schemes and full action space baselines to assess their impact on policy learning and performance. Our framework leverages a multimodal Proximal Policy Optimization agent that processes both semantic image sequences and scalar vehicle states. The proposed dynamic and relative strategies incorporate real-time action masking based on context and state transitions, preserving action consistency while eliminating invalid or suboptimal choices. Through comprehensive experiments across diverse driving routes, we show that action space reduction significantly improves training stability and policy performance. The dynamic and relative schemes, in particular, achieve a favorable balance between learning speed, control precision, and generalization. These findings highlight the importance of context-aware action space design for scalable and reliable RL in autonomous driving tasks.
Figures
Forward citations
Cited by 2 Pith papers
-
Optimality-Based Control Space Reduction for Infinite-Dimensional Control Spaces
A state-space reduction in linear-quadratic parabolic optimal control automatically induces an equivalent control-space reduction, with certified error bounds and a convergent adaptive algorithm.
-
A Review of Learning-Based Motion Planning: Toward a Data-Driven Optimal Control Approach
A position/review paper argues data-driven model predictive control is the best route to safe, adaptive, human-like autonomous-driving motion planning, but provides no new derivation or experiment.
Reference graph
Works this paper leans on
-
[1]
MuJoCo: A physics engine for model-based control | IEEE conference publication | IEEE xplore
-
[2]
Common- road: Composable benchmarks for motion planning on roads
Matthias Althoff, Markus Koschi, and Stefanie Manzinger. Common- road: Composable benchmarks for motion planning on roads. In 2017 IEEE Intelligent Vehicles Symposium (IV) , pages 719–726. IEEE, 2017
work page 2017
-
[3]
Reinforced Cur- riculum Learning For Autonomous Driving In Carla
Luca Anzalone, Silvio Barra, and Michele Nappi. Reinforced Cur- riculum Learning For Autonomous Driving In Carla. In 2021 IEEE International Conference on Image Processing (ICIP) , pages 3318– 3322, September 2021
work page 2021
-
[4]
Traffic Signal Control Using Hybrid Action Space Deep Reinforcement Learning
Salah Bouktif, Abderraouf Cheniki, and Ali Ouni. Traffic Signal Control Using Hybrid Action Space Deep Reinforcement Learning. Sensors, 21(7):2302, January 2021
work page 2021
-
[5]
Greg Brockman, Vicki Cheung, Ludwig Pettersson, Jonas Schneider, John Schulman, Jie Tang, and Wojciech Zaremba. OpenAI gym
-
[6]
GRI: General Reinforced Imitation and its Application to Vision-Based Autonomous Driving, May 2022
Raphael Chekroun, Marin Toromanoff, Sascha Hornauer, and Fabien Moutarde. GRI: General Reinforced Imitation and its Application to Vision-Based Autonomous Driving, May 2022
work page 2022
-
[7]
Learning to drive from a world on rails, October 2021
Dian Chen, Vladlen Koltun, and Philipp Kr ¨ahenb¨uhl. Learning to drive from a world on rails, October 2021. Comment: Pa- per published in ICCV 2021(Oral); Code and data available at: https://dotchen.github.io/world on rails/
work page 2021
-
[8]
Context- Aware Meta-RL With Two-Stage Constrained Adaptation for Urban Driving
Qi Deng, Ruyang Li, Qifu Hu, Yaqian Zhao, and Rengang Li. Context- Aware Meta-RL With Two-Stage Constrained Adaptation for Urban Driving. IEEE Transactions on Vehicular Technology, 73(2):1567–1581, February 2024. Conference Name: IEEE Transactions on Vehicular Technology
work page 2024
Show all 56 references
-
[9]
Context - Enhanced Meta-Reinforcement Learning with Data-Reused Adaptation for Urban Autonomous Driving
Qi Deng, Yaqian Zhao, Rengang Li, Qifu Hu, Tiejun Liu, and Ruyang Li. Context - Enhanced Meta-Reinforcement Learning with Data-Reused Adaptation for Urban Autonomous Driving. In 2023 International Joint Conference on Neural Networks (IJCNN) , pages 1–8, June 2023. ISSN: 2161-4407
2023
-
[10]
CARLA: An Open Urban Driving Simulator
Alexey Dosovitskiy, German Ros, Felipe Codevilla, Antonio Lopez, and Vladlen Koltun. CARLA: An Open Urban Driving Simulator. In Proceedings of the 1st Annual Conference on Robot Learning , pages 1–16. PMLR, October 2017
2017
-
[11]
Deep reinforcement learning in large discrete action spaces
Gabriel Dulac-Arnold, Richard Evans, Hado van Hasselt, Peter Sunehag, Timothy Lillicrap, Jonathan Hunt, Timothy Mann, Theophane Weber, Thomas Degris, and Ben Coppin. Deep reinforcement learning in large discrete action spaces
-
[12]
DQN-based Reinforcement Learning for Vehicle Control of Autonomous Vehicles Interacting With Pedestrians
Badr Ben Elallid, Nabil Benamar, Nabil Mrani, and Tajjeeddine Rachidi. DQN-based Reinforcement Learning for Vehicle Control of Autonomous Vehicles Interacting With Pedestrians. In 2022 International Conference on Innovation and Intelligence for Informatics, Computing, and Tech...
2022
-
[13]
Growing Action Spaces, June 2019
Gregory Farquhar, Laura Gustafson, Zeming Lin, Shimon Whiteson, Nicolas Usunier, and Gabriel Synnaeve. Growing Action Spaces, June 2019
2019
-
[14]
Intersection decision making for autonomous vehicles based on improved PPO algorithm
Dong Guo, Shoulin He, and Shouwen Ji. Intersection decision making for autonomous vehicles based on improved PPO algorithm. IET Intelligent Transport Systems , 18(S1):2921–2938, 2024. eprint: https://onlinelibrary.wiley.com/doi/pdf/10.1049/itr2.12593
2024 doi
-
[15]
Interactive fiction games: A colossal adventure
Matthew Hausknecht, Prithviraj Ammanabrolu, Marc-Alexandre C ˆot´e, and Xingdi Yuan. Interactive fiction games: A colossal adventure. 34(5):7903–7910. Number: 05
-
[16]
Le, and Hartwig Adam
Andrew Howard, Mark Sandler, Grace Chu, Liang-Chieh Chen, Bo Chen, Mingxing Tan, Weijun Wang, Yukun Zhu, Ruoming Pang, Vijay Vasudevan, Quoc V . Le, and Hartwig Adam. Searching for mobilenetv3. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) ,...
2019
-
[17]
A closer look at invalid action masking in policy gradient algorithms
Shengyi Huang and Santiago Onta ˜n´on. A closer look at invalid action masking in policy gradient algorithms. 35
-
[18]
Safe Reinforcement Learning on Autonomous Vehicles
David Isele, Alireza Nakhaei, and Kikuo Fujimura. Safe Reinforcement Learning on Autonomous Vehicles. In 2018 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages 1–6, October 2018. ISSN: 2153-0866
2018
-
[19]
Action Space Shaping in Deep Reinforcement Learning, May 2020
Anssi Kanervisto, Christian Scheller, and Ville Hautam ¨aki. Action Space Shaping in Deep Reinforcement Learning, May 2020. Comment: To appear in IEEE Conference on Games 2020. Experiment code is available at https://github.com/Miffyli/rl-action-space-shaping
2020
-
[20]
ViZDoom: A doom-based AI research platform for visual reinforcement learning
Michał Kempka, Marek Wydmuch, Grzegorz Runc, Jakub Toczek, and Wojciech Ja ´skowski. ViZDoom: A doom-based AI research platform for visual reinforcement learning. In 2016 IEEE Conference on Compu- tational Intelligence and Games (CIG) , pages 1–8. ISSN: 2325-4289
2016
-
[21]
Learning to Drive in a Day
Alex Kendall, Jeffrey Hawke, David Janz, Przemyslaw Mazur, Daniele Reda, John-Mark Allen, Vinh-Dieu Lam, Alex Bewley, and Amar Shah. Learning to Drive in a Day. In 2019 International Conference on Robotics and Automation (ICRA) , pages 8248–8254, May 2019
2019
-
[22]
Al Sallab, Senthil Yogamani, and Patrick P ´erez
B Ravi Kiran, Ibrahim Sobh, Victor Talpaert, Patrick Mannion, Ahmad A. Al Sallab, Senthil Yogamani, and Patrick P ´erez. Deep reinforcement learning for autonomous driving: A survey. 23(6):4909–4926
-
[23]
The highD dataset: A drone dataset of naturalistic vehicle trajectories on german highways for validation of highly automated driving systems
Robert Krajewski, Julian Bock, Laurent Kloeker, and Lutz Eckstein. The highD dataset: A drone dataset of naturalistic vehicle trajectories on german highways for validation of highly automated driving systems
-
[24]
Safe Rein- forcement Learning for Autonomous Lane Changing Using Set-Based Prediction
Hanna Krasowski, Xiao Wang, and Matthias Althoff. Safe Rein- forcement Learning for Autonomous Lane Changing Using Set-Based Prediction. In 2020 IEEE 23rd International Conference on Intelligent Transportation Systems (ITSC) , pages 1–7, September 2020
2020
-
[25]
HyAR: Addressing discrete- continuous action reinforcement learning via hybrid action representa- tion
Boyan Li, Hongyao Tang, Yan Zheng, Jianye Hao, Pengyi Li, Zhen Wang, Zhaopeng Meng, and Li Wang. HyAR: Addressing discrete- continuous action reinforcement learning via hybrid action representa- tion
-
[26]
Think2Drive: Efficient Reinforcement Learning by Thinking in Latent World Model for Quasi-Realistic Autonomous Driving (in CARLA-v2), July 2024
Qifeng Li, Xiaosong Jia, Shaobo Wang, and Junchi Yan. Think2Drive: Efficient Reinforcement Learning by Thinking in Latent World Model for Quasi-Realistic Autonomous Driving (in CARLA-v2), July 2024. Comment: Accepted by ECCV 2024
2024
-
[27]
MetaDrive: Composing diverse driving scenarios for generalizable reinforcement learning
Quanyi Li, Zhenghao Peng, Lan Feng, Qihang Zhang, Zhenghai Xue, and Bolei Zhou. MetaDrive: Composing diverse driving scenarios for generalizable reinforcement learning. 45(3):3461–3475
-
[28]
Achieving sample and computational efficient reinforcement learning by action space reduction via grouping
Yining Li, Peizhong Ju, and Ness Shroff. Achieving sample and computational efficient reinforcement learning by action space reduction via grouping
-
[29]
Achieving Sample and Com- putational Efficient Reinforcement Learning by Action Space Reduction via Grouping, June 2023
Yining Li, Peizhong Ju, and Ness Shroff. Achieving Sample and Com- putational Efficient Reinforcement Learning by Action Space Reduction via Grouping, June 2023
2023
-
[30]
Lillicrap, Jonathan J
Timothy P. Lillicrap, Jonathan J. Hunt, Alexander Pritzel, Nicolas Heess, Tom Erez, Yuval Tassa, David Silver, and Daan Wierstra. Continuous control with deep reinforcement learning. arXiv preprint arXiv:1509.02971, 2015
2015 arXiv
-
[31]
Exact Reduction of Huge Action Spaces in General Reinforcement Learning, December 2020
Sultan Javed Majeed and Marcus Hutter. Exact Reduction of Huge Action Spaces in General Reinforcement Learning, December 2020. Comment: A variant of this paper was presented at the AAAI-2021 proceedings
2020
-
[32]
Webotstm: Professional mobile robot simulation, 2004
Olivier Michel. Webotstm: Professional mobile robot simulation, 2004
2004
-
[33]
Safe and Psychologically Pleas- ant Traffic Signal Control with Reinforcement Learning using Action Masking
Arthur M ¨uller and Matthia Sabatelli. Safe and Psychologically Pleas- ant Traffic Signal Control with Reinforcement Learning using Action Masking. In 2022 IEEE 25th International Conference on Intelligent Transportation Systems (ITSC) , pages 951–958, October 2022
2022
-
[34]
Bergasa, Carlos G ´omez-Hu´elamo, Rodrigo Guti ´errez, and Alejandro D ´ıaz-D´ıaz
´Oscar P ´erez-Gil, Rafael Barea, Elena L ´opez-Guill´en, Luis M. Bergasa, Carlos G ´omez-Hu´elamo, Rodrigo Guti ´errez, and Alejandro D ´ıaz-D´ıaz. Deep reinforcement learning based control for Autonomous Vehicles in CARLA. Multimedia Tools and Applications, 81(3):3553–3576, ...
2022
-
[35]
Stable-baselines3: Reliable reinforcement learning implementations
Antonin Raffin, Ashley Hill, Adam Gleave, Anssi Kanervisto, Max- imilian Ernestus, and Noah Dormann. Stable-baselines3: Reliable reinforcement learning implementations. Journal of Machine Learning Research, 22(268):1–8, 2021
2021
-
[36]
Fuzzy Action-Masked Reinforcement Learning Behav- ior Planning for Highly Automated Driving
Thomas Rudolf, Mingxiao Gao, Tobias Sch ¨urmann, Stefan Schwab, and S¨oren Hohmann. Fuzzy Action-Masked Reinforcement Learning Behav- ior Planning for Highly Automated Driving. In 2022 8th International Conference on Control, Automation and Robotics (ICCAR) , pages 264– 270, A...
2022
-
[37]
Proximal Policy Optimization Algorithms, August 2017
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov. Proximal Policy Optimization Algorithms, August 2017
2017
-
[38]
Autonomous driving in CARLA using deep reinforcement learning
Idrees Shaikh. Autonomous driving in CARLA using deep reinforcement learning. https://github.com/idreesshaikh/ Autonomous-Driving-in-Carla-using-Deep-Reinforcement-Learning,
-
[39]
Excluding the irrelevant: Focusing reinforcement learning through continuous action masking
Roland Stolz, Hanna Krasowski, Jakob Thumm, Michael Eichelbeck, Philipp Gassert, and Matthias Althoff. Excluding the irrelevant: Focusing reinforcement learning through continuous action masking
-
[40]
Dynamic Action Space Handling Method for Reinforcement Learning Models
Sangchul Woo and Yunsick Sung. Dynamic Action Space Handling Method for Reinforcement Learning Models. Journal of Information Processing Systems , 16(5):1223–1230, October 2020. [TLDR] This study proposes a method of reinforcement learning in which the action space is dynamica...
2020
-
[41]
Action branch- ing architectures for deep reinforcement learning
Arash Tavakoli, Fabio Pardo, and Petar Kormushev. Action branch- ing architectures for deep reinforcement learning. arXiv preprint arXiv:1711.08946, 2017
2017 arXiv
-
[42]
Action Branch- ing Architectures for Deep Reinforcement Learning, January 2019
Arash Tavakoli, Fabio Pardo, and Petar Kormushev. Action Branch- ing Architectures for Deep Reinforcement Learning, January 2019. Comment: AAAI 2018, NIPS 2017 Deep RL Symposium, code: https://github.com/atavakol/action-branching-agents
2019
-
[43]
End-to-End Model-Free Reinforcement Learning for Urban Driving using Implicit Affordances, March 2020
Marin Toromanoff, Emilie Wirbel, and Fabien Moutarde. End-to-End Model-Free Reinforcement Learning for Urban Driving using Implicit Affordances, March 2020. Comment: Accepted at main conference of CVPR 2020
2020
-
[44]
Deep Reinforcement Learning With Action Masking for Differential-Drive Robot Navigation Using Low-Cost Sen- sors
Konstantinos Tsampazis, Manos Kirtas, Pavlos Tosidis, Nikolaos Pas- salis, and Anastasios Tefas. Deep Reinforcement Learning With Action Masking for Differential-Drive Robot Navigation Using Low-Cost Sen- sors. In 2023 IEEE 33rd International Workshop on Machine Learning for S...
2023
-
[45]
Pure-past action masking
Giovanni Varricchione, Natasha Alechina, Mehdi Dastani, Giuseppe De Giacomo, Brian Logan, and Giuseppe Perelli. Pure-past action masking. 38(19):21646–21655. Number: 19
-
[46]
StarCraft II: A new challenge for reinforcement learning
Oriol Vinyals, Timo Ewalds, Sergey Bartunov, Petko Georgiev, Alexan- der Sasha Vezhnevets, Michelle Yeo, Alireza Makhzani, Heinrich K¨uttler, John Agapiou, Julian Schrittwieser, John Quan, Stephen Gaffney, Stig Petersen, Karen Simonyan, Tom Schaul, Hado van Hasselt, David Silv...
-
[47]
A Versatile and Efficient Reinforcement Learning Framework for Autonomous Driving, March 2022
Guan Wang, Haoyi Niu, Desheng Zhu, Jianming Hu, Xianyuan Zhan, and Guyue Zhou. A Versatile and Efficient Reinforcement Learning Framework for Autonomous Driving, March 2022. Comment: 8 pages, 6 figures
2022
-
[48]
Continuous control for automated lane change behavior based on deep deterministic policy gradient algorithm
Pin Wang, Hanhan Li, and Ching-Yao Chan. Continuous control for automated lane change behavior based on deep deterministic policy gradient algorithm. In 2019 IEEE Intelligent Vehicles Symposium (IV) , pages 1454–1460. ISSN: 2642-7214
2019
-
[49]
Learning State-Specific Action Masks for Reinforcement Learning
Ziyi Wang, Xinran Li, Luoyang Sun, Haifeng Zhang, Hualin Liu, and Jun Wang. Learning State-Specific Action Masks for Reinforcement Learning. Algorithms, 17(2):60, February 2024
2024
-
[50]
Reinforce- ment Learning-based Autonomous Parking with Expert Demonstrations
Yao Wu, Lucai Wang, Xiao Lu, Yue Wu, and Haojun Zhang. Reinforce- ment Learning-based Autonomous Parking with Expert Demonstrations. In 2023 7th CAA International Conference on Vehicular Control and Intelligence (CVCI), pages 1–6, October 2023
2023
-
[51]
Fuzzy sets, fuzzy logic, and fuzzy systems: selected papers , volume 6
Lotfi Asker Zadeh, George J Klir, and Bo Yuan. Fuzzy sets, fuzzy logic, and fuzzy systems: selected papers , volume 6. World scientific, 1996
1996
-
[52]
Mankowitz, and Shie Mannor
Tom Zahavy, Matan Haroush, Nadav Merlis, Daniel J. Mankowitz, and Shie Mannor. Learn What Not to Learn: Action Elimination with Deep Reinforcement Learning, February 2019
2019
-
[53]
Lexicographic Actor-Critic Deep Reinforcement Learning for Urban Autonomous Driving
Hengrui Zhang, Youfang Lin, Sheng Han, and Kai Lv. Lexicographic Actor-Critic Deep Reinforcement Learning for Urban Autonomous Driving. IEEE Transactions on Vehicular Technology , 72(4):4308– 4319, April 2023. Conference Name: IEEE Transactions on Vehicular Technology
2023
-
[54]
End-to-End Urban Driving by Imitating a Reinforcement Learning Coach, October 2021
Zhejun Zhang, Alexander Liniger, Dengxin Dai, Fisher Yu, and Luc Van Gool. End-to-End Urban Driving by Imitating a Reinforcement Learning Coach, October 2021. Comment: Published at ICCV 2021
2021
-
[55]
SMARTS: An open-source scalable multi-agent RL training school for autonomous driving
Ming Zhou, Jun Luo, Julian Villella, Yaodong Yang, David Rusu, Jiayu Miao, Weinan Zhang, Montgomery Alban, Iman Fadakar, Zheng Chen, Chongxi Huang, Ying Wen, Kimia Hassanzadeh, Daniel Graves, Zheng- bang Zhu, Yihan Ni, Nhat Nguyen, Mohamed Elsayed, Haitham Am- mar, Alexander C...
2020
-
[2023]
GitHub repository, commit 3f1c7d9, accessed 23 Apr 2025
2025
Reviewed August 6, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.