REVIEW 1 major objections 7 minor 5 cited by
CaRL: Learning Scalable Planning Policies with Simple Rewards
T0 review · 1 major / 7 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Replacing the usual multi-term rewards with a single route-completion term lets a standard RL algorithm scale to 300M samples and reach 64 Driving Score, far above prior RL planners.
desk verdict A genuinely strong RL-for-driving paper whose headline scalability claim is undermined by a missing control; the method and nuPlan results are still worth taking seriously. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the single-term reward $r_t = RC_t \left(\prod p_t\right) - T$, where $RC_t$ is the fraction of the route completed during the current simulator step, $p_t \in [0,1]$ are soft penalty factors that multiplicatively shrink the reward while the vehicle violates constraints such as the speed limit, lane-centre distance, time-to-collision, or comfort bounds, and $T$ is a terminal penalty applied at episode end for hard infractions ($T=1$ for collisions and red-light violations, $0$ otherwise). Major infractions terminate the episode, so violating them forfeits all future route completion. The paper motivates three properties of this design: the total reward obtainable is finite (only 100 route-completion points exist), no rule-based planner contributes to any reward term, and the reward's global optimum is the same as the global optimum of the Driving Score metric. The supporting machinery is the scaling stack built around it: Proximal Policy Optimization (PPO) hyperparameters in the Atari style that reduce off-policy gradient steps from 959 to 15, an asynchronous collection loop, DD-PPO (distributed data-parallel PPO) for scaling, and a bird's-eye-view input with a much wider field of view than prior RL planners.
What would settle it
Evaluate the 300M-sample CARLA model on the official CARLA leaderboard 2.0 validation routes (the hand-authored 20-route set on town 13) or on a set of held-out towns never seen in training. If the Driving Score collapses relative to the 64 achieved on longest6 v2—or if a variant trained only on the original hand-crafted scenario definitions matches it—then the auto-generated training distribution, rather than the reward design, is carrying the result. A second, cheaper check replays the mini-batch comparison of Table 3 across seeds: the paper's core claim stands only if the complex reward consistently collapses at mini-batch 1024 while the route-completion reward consistently improves.
Extended reading notes
Core claim
The paper's central claim is that a driving reward consisting only of route completion, with infractions encoded as episode termination or multiplicative penalties, changes how PPO behaves under scaling: the same algorithm that collapses to 2±2 Driving Score at mini-batch size 1024 with a complex multi-term reward improves to 38±3 with the simple reward, and reaches 64±2 at mini-batch 16384 with 300 million samples. This is presented as evidence that previous RL planners were limited by reward design rather than by data or algorithm: the complex reward's handcrafted terms introduce local minima and rule-based upper bounds on performance, whereas the route-completion reward has a finite reward budget (only 100 route-completion points exist), contains no rule-based planner, and shares its global optimum with the evaluation metric. On nuPlan the identical recipe, adapted only by adding a survival bonus for the simulator's fixed-length episodes, achieves 91.3 closed-loop score in non-reactive and 90.6 in reactive traffic on the Val14 benchmark, making it the best learning-based method while running an order of magnitude faster at inference than prior work.
Load-bearing premise
The load-bearing premise is that the training routes and scenarios generated automatically by rejection sampling and heuristics, rather than by human labelling, are diverse and representative enough that a policy trained on them transfers to the fixed evaluation routes of longest6 v2; if those generated scenarios are easier or biased relative to the test routes, the reported Driving Scores would overstate the policy's real planning ability.
Editorial extensions
If this is right
- RL planners for driving can be scaled with data the way imitation learners are: under the simple reward, raising the sample count from 10 million to 300 million with a larger mini-batch increases Driving Score by 33 points, the largest single improvement reported in this line of work.
- Rule-based planners can be dropped from the reward pipeline entirely: the resulting policy outperforms every prior RL planner on CARLA and every learning-based planner on nuPlan without any scenario-specific rules.
- Large mini-batch sizes become an asset rather than a liability for on-policy RL in driving, so distributed data-parallel collection (DD-PPO) is the right scaling regime for this task.
- The same reward and hyperparameters transfer across simulators with only a survival-bonus adaptation, suggesting the design exploits a general property of goal-directed driving rather than CARLA-specific quirks.
- Because the reward's global optimum equals the evaluation metric's, further gains should come from more data and model capacity rather than reward re-tuning; the remaining gap to the rule-based PDM-Lite is attributed to consistency in solving safety-critical scenarios, not to reward misalignment.
Reading between the lines
- The property that likely carries the scaling result is the finiteness of the reward budget: with only 100 route-completion points available, no stationary behaviour can harvest reward forever, which is exactly the loophole the paper documents in the complex reward (waiting at green lights); any long-horizon task with a natural 'fraction of goal achieved' signal—navigation, manipulation, search—may
- The paper's explanation for the complex-reward collapse—that larger mini-batches smooth the gradient and push early optimization into the broadest local optimum (not driving)—is testable directly: measuring per-term gradient correlations or loss-landscape curvature as a function of mini-batch size would confirm or refute it without any new hardware.
- Because training and evaluation share the same towns (level 4), the natural next claim to test is generalization: the reward and inputs carry no town-specific rules, so the recipe is a plausible starting point for zero-shot transfer to unseen towns, which the paper explicitly leaves open.
- The nuPlan survival bonus is a workaround for episodes of fixed length; in open-ended operation with no defined route end, route completion alone would stop giving signal once the goal is reached, so any deployment beyond benchmarks would need a terminal-state design or an explicit 'keep driving safely' term.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes CaRL, a privileged RL planning method for CARLA and nuPlan trained with PPO and a sparse reward that rewards route completion, terminates episodes on hard infractions, and applies multiplicative soft penalties for constraint violations. The authors report that this reward, unlike a shaped Roach-style reward, allows PPO to benefit from larger mini-batch sizes, enabling scaling to 300M samples in CARLA and 500M/1B samples in nuPlan on a single 8-GPU node. They report 64 DS on the longest6 v2 benchmark, outperforming prior RL planners by a large margin, and 91.3/90.6 CLS on nuPlan Val14, outperforming prior learning-based planners. The paper includes ablations with multiple seeds, a released codebase, and reimplementations of baselines.
Significance. If the central scalability claim holds, this is a noteworthy result for RL-based driving: it challenges the prevailing shaped-reward paradigm, demonstrates that a metric-aligned sparse reward can be optimized at scale with PPO, and provides a reproducible recipe for CARLA and nuPlan. The engineering contributions, including the RL-optimized leaderboard code, AC-PPO, and large-scale training, are also valuable. The paper's reproducibility practices are strong: code is released, experiments average multiple seeds, baselines are reimplemented, and evaluation uses the official leaderboard code. However, the key comparison used to attribute the scalability advantage to reward design is confounded, so the significance of the conceptual claim is currently uncertain pending a matched-update experiment.
major comments (1)
- [Section 3, Table 3; Section 3.1] Section 3.1 states that the mini-batch 1024 runs were obtained by 'performing 4x more simulator steps per PPO iteration with 4x fewer PPO iterations' while holding total samples fixed. This means the comparison in Table 3 changes three variables simultaneously: mini-batch size (256 to 1024), rollout length per PPO iteration (128 to 512), and total number of optimizer updates (reduced by a factor of 4). Longer rollouts directly improve return estimation for a sparse reward, and fewer updates can explain the collapse of the Roach reward (34±7 to 2±2) as undertraining rather than a reward-induced local minimum. The text's attribution of the collapse to 'larger mini-batch sizes smooth the optimization' (Section 3) is therefore not supported by the evidence presented. Please run an ablation that varies mini-batch size while holding rollout length and total gradient updates fixed, e.g., by increasing the number of parallel environments instead of the rollout length and keeping the number of PPO iterations constant; report learning curves and total update counts. If the complex reward recovers with a matched update budget, the abstract's claim that complex rewards 'limit scalability' should be revised to a compute-budget effect rather than a reward-design effect.
minor comments (7)
- [Section 3, Eq. (1), principle (3)] The statement that 'the global optimum of the reward is the same as the global optimum of the metric' is an overstatement: the reward includes soft penalties for speeding, TTC, and comfort that are not part of the CARLA DS metric, so a policy that maximizes DS but violates these soft constraints would receive suboptimal reward. Please soften this to 'closely aligned' and state the residual differences.
- [Section 3.1, Table 3] Please state in the table caption or text the rollout length and total number of gradient updates used in each column; currently the reader cannot disentangle mini-batch size from these other changes.
- [Abstract and Introduction] The phrase 'single intuitive reward term' is somewhat misleading because Eq. (1) includes several multiplicative soft penalties and terminal penalties; consider describing it as a single primary reward term with constraint penalties.
- [Table 4 and Section 3.1] The 300M model is averaged over three seeds while other experiments use five; please justify the reduced number of seeds for the main result.
- [Section 4.2, Table 6] The claim of being the best learning-based method on nuPlan should be qualified: PLUTO (92.6 NR) is excluded as a hybrid IL+rule method, and the main-text comparison does not include it in the table. Please make the exclusion criterion explicit and consider reporting PLUTO in Table 6.
- [Appendix D.2] The automatically generated training routes and scenarios are acknowledged to be noisy and unbalanced; adding basic statistics, such as number of routes, scenario type distribution, and rejection rates, would help assess the generalization evidence.
- [Appendix E.1.2] Because Think2Drive's code is not public and key hyperparameters were inferred, the reported 7 DS result may be sensitive to these choices; please provide a short sensitivity analysis or release the reproduction to strengthen the comparison.
Circularity Check
No significant circularity: reward–metric alignment is a stated design choice, the scaling claim is empirical, and the nuPlan-tuned variant is disclosed rather than presented as a prediction.
full rationale
The central derivation—that a route-completion reward with terminal/soft penalties allows PPO to scale to large mini-batches and 300M samples—is an empirical result, not a consequence of definition. Section 3 explicitly states design principle (3) that the reward's global optimum matches the metric's optimum; this is a deliberate objective alignment (standard reward design) rather than a hidden equivalence, and the paper confirms the reward 'closely aligns with the DS metric' while adding extra penalties. The reported CARLA DS is not a fitted parameter: no aspect of the reward or architecture is tuned to longest6 v2 scores beyond the stated reward form, and comparisons to Roach/Think2Drive/PDM-Lite are independent reproductions. The nuPlan 'CaRL (nuPlan tuned)' variant in Appendix E.2 does take the CLS metric's own 5:5:4:2 weights and add TTC/comfort/speed terms to route completion, so its improved CLS is partly by construction; however, the paper labels it 'nuPlan tuned', discloses the exact weight copying, warns that the reward is not recommended generally, and does not use it for the headline 91.3/90.6 claims. Self-citations ([17] for reward–metric alignment, [51] for benchmarking practice) are not load-bearing: the alignment principle is domain-general and the benchmark results stand on the paper's own experiments. The Table 3 mini-batch comparison confounds mini-batch size with number of PPO updates (4x fewer iterations at 1024), but this is a validity/confound issue rather than circularity, and it is not a case where a fitted input is renamed as a prediction. Overall, no derivation step reduces to its own input by construction.
Assumptions & free parameters
free parameters (8)
- survival_ratio_s =
0.6 (nuPlan)
- terminal_penalty_T =
1.0 for collision/red light, 0.0 otherwise
- soft_penalty_factors =
pt = 0 (outside lanes), 0.5 (TTC), 1 - 0.5*#/6 (comfort)
- comfort_thresholds =
Table 16 (e.g., longitudinal accel (-20,10) m/s^2 in CARLA)
- route_deviation_threshold =
30 m
- blocked_threshold =
90 s at <0.1 m/s
- action_scaling_nuPlan =
[-3.2, 2.4] m/s^2, [-0.84, 0.84] rad
- train_frequency =
10 Hz
assumptions (5)
- domain assumption The CARLA and nuPlan simulators provide sufficiently faithful models of urban driving such that policies optimized in them transfer to the evaluation protocol.
- domain assumption The BEV semantic observation, with the modifications in Section B.5, contains sufficient information to drive safely.
- domain assumption PPO with the Atari hyperparameters is a sufficiently capable optimizer for the simple reward at scale.
- domain assumption The global optimum of the proposed reward (Eq. 1) aligns with the global optimum of the driving metric (DS/CLS).
- domain assumption The automatically generated training routes and scenarios are representative of the evaluation distribution.
Cite this review
Pith. "Pith review of CaRL: Learning Scalable Planning Policies with Simple Rewards." pith.science (2026). https://pith.science/paper/XK2QDROZ
@misc{pith2026250417838,
author = {Pith},
title = {Pith review of: CaRL: Learning Scalable Planning Policies with Simple Rewards},
year = {2026},
howpublished = {\url{https://pith.science/paper/XK2QDROZ}},
note = {Machine review of arXiv:2504.17838}
}
read the original abstract
We investigate reinforcement learning (RL) for privileged planning in autonomous driving. State-of-the-art approaches for this task are rule-based, but these methods do not scale to the long tail. RL, on the other hand, is scalable and does not suffer from compounding errors like imitation learning. Contemporary RL approaches for driving use complex shaped rewards that sum multiple individual rewards, \eg~progress, position, or orientation rewards. We show that PPO fails to optimize a popular version of these rewards when the mini-batch size is increased, which limits the scalability of these approaches. Instead, we propose a new reward design based primarily on optimizing a single intuitive reward term: route completion. Infractions are penalized by terminating the episode or multiplicatively reducing route completion. We find that PPO scales well with higher mini-batch sizes when trained with our simple reward, even improving performance. Training with large mini-batch sizes enables efficient scaling via distributed data parallelism. We scale PPO to 300M samples in CARLA and 500M samples in nuPlan with a single 8-GPU node. The resulting model achieves 64 DS on the CARLA longest6 v2 benchmark, outperforming other RL methods with more complex rewards by a large margin. Requiring only minimal adaptations from its use in CARLA, the same method is the best learning-based approach on nuPlan. It scores 91.3 in non-reactive and 90.6 in reactive traffic on the Val14 benchmark while being an order of magnitude faster than prior work.
Figures
Figures from the paper (7 more)
Forward citations
Cited by 5 Pith papers
-
On Data Thinning for Model Validation in Small Area Estimation
Thinned-data MSE for small-area models is unbiased for a risk that systematically differs from full-data risk; under Fay-Herriot the gap is closed-form in the model's shrinkage, and the thinning fraction faces a sharp...
-
Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning
TPSP trains a scene-perturbation network using threat differences between original and perturbed rollouts to create safety-critical experiences for an online RL driving policy.
-
Zero-Human Demonstration End-to-end Autonomous Driving with Trajectory Scorer
A reward-only offline RL method for trajectory planning in end-to-end autonomous driving achieves state-of-the-art on Navhard and competitive closed-loop HUGSIM performance without imitation learning.
-
CLEAR: Closed-Loop Reinforcement Learning at Scale for End-to-End Autonomous Driving
Residual waypoint RL around a frozen VLA prior, scaled via heterogeneous CARLA/H100 infrastructure, raises closed-loop driving score and success rate on longest6 v2 and Bench2Drive.
-
End-to-End Crop Row Navigation via LiDAR-Based Deep Reinforcement Learning
Raw 3D LiDAR, compressed into flattened voxel maps, trains a reinforcement learning policy that reliably follows straight crop rows in simulation and degrades on curvier rows.
Reference graph
Works this paper leans on
-
[1]
Treiber, A
M. Treiber, A. Hennecke, and D. Helbing. Congested traffic states in empirical observations and micro- scopic simulations. Physical review E, 2000
2000
-
[2]
Thrun, M
S. Thrun, M. Montemerlo, H. Dahlkamp, D. Stavens, A. Aron, J. Diebel, P. Fong, J. Gale, M. Halpenny, G. Hoffmann, K. Lau, C. M. Oakley, M. Palatucci, V . R. Pratt, P. Stang, S. Strohband, C. Dupont, L. Jendrossek, C. Koelen, C. Markey, C. Rummel, J. van Niekerk, E. Jensen, P. Alessandrini, G. R. Bradski, B. Davies, S. Ettinger, A. Kaehler, A. V . Nefian, ...
2006
-
[3]
B. Jaeger. Expert drivers for autonomous driving. Master’s thesis, University of Tübingen, 2021
2021
-
[4]
Dauner, M
D. Dauner, M. Hallgarten, A. Geiger, and K. Chitta. Parting with misconceptions about learning-based vehicle motion planning. In Proc. Conf. on Robot Learning (CoRL), 2023
2023
-
[5]
C. Sima, K. Renz, K. Chitta, L. Chen, H. Zhang, C. Xie, J. Beißwenger, P. Luo, A. Geiger, and H. Li. Drivelm: Driving with graph visual question answering. In Proc. of the European Conf. on Computer Vision (ECCV), 2024
2024
-
[6]
K. Renz, K. Chitta, O.-B. Mercea, A. S. Koepke, Z. Akata, and A. Geiger. Plant: Explainable planning transformers via object-level representations. In Proc. Conf. on Robot Learning (CoRL), 2022. 9
2022
-
[7]
Hallgarten, M
M. Hallgarten, M. Stoll, and A. Zell. From prediction to planning with goal conditioned lane graph traversals. In Proc. IEEE Conf. on Intelligent Transportation Systems (ITSC), 2023
2023
-
[8]
Cheng, Y
J. Cheng, Y . Chen, X. Mei, B. Yang, B. Li, and M. Liu. Rethinking imitation-based planners for au- tonomous driving. In Proc. IEEE International Conf. on Robotics and Automation (ICRA), 2024
2024
Show all 110 references
-
[9]
Huang, H
Z. Huang, H. Liu, J. Wu, and C. Lv. Differentiable integrated motion prediction and planning with learnable cost function for autonomous driving. IEEE Trans. Neural Networks Learn. Syst., 2024
2024
-
[10]
Q. Sun, H. Wang, J. Zhan, F. Nie, X. Wen, L. Xu, K. Zhan, P. Jia, X. Lang, and H. Zhao. Generalizing motion planners with mixture of experts for autonomous driving. arXiv.org, 2410.15774, 2024
2024 arXiv
-
[11]
J. Guo, M. Feng, P. Zhu, C. Li, and J. Pu. Rethinking closed-loop planning framework for imitation- based model integrating prediction and planning. arXiv.org, 2407.05376, 2024
2024 arXiv
-
[12]
Cheng, Y
J. Cheng, Y . Chen, and Q. Chen. PLUTO: pushing the limit of imitation learning-based planning for autonomous driving. arXiv.org, 2404.14327, 2024
2024 arXiv
-
[13]
X. Chen, J. Yan, W. Liao, T. He, and P. Peng. Int2planner: An intention-based multi-modal motion planner for integrated prediction and planning. arXiv.org, 2501.12799, 2025
2025 arXiv
-
[14]
Zhang, A
Z. Zhang, A. Liniger, D. Dai, F. Yu, and L. Van Gool. End-to-end urban driving by imitating a reinforce- ment learning coach. In Proc. of the IEEE International Conf. on Computer Vision (ICCV), 2021
2021
-
[15]
Zhang, R
C. Zhang, R. Guo, W. Zeng, Y . Xiong, B. Dai, R. Hu, M. Ren, and R. Urtasun. Rethinking closed-loop training for autonomous driving. In Proc. of the European Conf. on Computer Vision (ECCV), 2022
2022
-
[16]
R. S. Sutton and A. G. Barto. Reinforcement learning: An introduction. MIT press, 2018
2018
-
[17]
Jaeger and A
B. Jaeger and A. Geiger. An invitation to deep reinforcement learning. Foundations and Trends in Optimization, 2024
2024
-
[18]
W. B. Knox, A. Allievi, H. Banzhaf, F. Schmitt, and P. Stone. Reward (mis)design for autonomous driving. Artificial Intelligence (AI), 2023
2023
-
[19]
Kendall, J
A. Kendall, J. Hawke, D. Janz, P. Mazur, D. Reda, J. Allen, V . Lam, A. Bewley, and A. Shah. Learning to drive in a day. In Proc. IEEE International Conf. on Robotics and Automation (ICRA), 2019
2019
-
[20]
Wijmans, A
E. Wijmans, A. Kadian, A. Morcos, S. Lee, I. Essa, D. Parikh, M. Savva, and D. Batra. DD-PPO: learning near-perfect pointgoal navigators from 2.5 billion frames. In Proc. of the International Conf. on Learning Representations (ICLR), 2020
2020
-
[21]
Fuchs, Y
F. Fuchs, Y . Song, E. Kaufmann, D. Scaramuzza, and P. Dürr. Super-human performance in gran turismo sport using deep reinforcement learning. IEEE Robotics and Automation Letters (RA-L), 2021
2021
-
[22]
K. Zeng, Z. Zhang, K. Ehsani, R. Hendrix, J. Salvador, A. Herrasti, R. B. Girshick, A. Kembhavi, and L. Weihs. Poliformer: Scaling on-policy RL with transformers results in masterful navigators. In Proc. Conf. on Robot Learning (CoRL), volume 2406.20083, 2024
2024 arXiv
-
[23]
P. R. Wurman, S. Barrett, K. Kawamoto, J. MacGlashan, K. Subramanian, T. J. Walsh, R. Capobianco, A. Devlic, F. Eckert, F. Fuchs, L. Gilpin, P. Khandelwal, V . Kompella, H. Lin, P. MacAlpine, D. Oller, T. Seno, C. Sherstan, M. D. Thomure, H. Aghabozorgi, L. Barrett, R. Douglas...
2022
-
[24]
Schulman, F
J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algo- rithms. arXiv.org, 1707.06347, 2017
2017 arXiv
-
[25]
Toromanoff, E
M. Toromanoff, E. Wirbel, and F. Moutarde. End-to-end model-free reinforcement learning for urban driving using implicit affordances. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2020
2020
-
[26]
Chekroun, M
R. Chekroun, M. Toromanoff, S. Hornauer, and F. Moutarde. GRI: general reinforced imitation and its application to vision-based autonomous driving. Robotics, 2023
2023
-
[27]
Q. Li, X. Jia, S. Wang, and J. Yan. Think2drive: Efficient reinforcement learning by thinking with latent world model for autonomous driving (in CARLA-V2). In Proc. of the European Conf. on Computer Vision (ECCV), 2024. 10
2024
-
[28]
Dosovitskiy, G
A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V . Koltun. CARLA: An open urban driving simu- lator. In Proc. Conf. on Robot Learning (CoRL), 2017
2017
-
[29]
C. team. Carla autonomous driving leaderboard 2.0. https://leaderboard.carla.org/, 2022
2022
-
[30]
Chitta, A
K. Chitta, A. Prakash, B. Jaeger, Z. Yu, K. Renz, and A. Geiger. Transfuser: Imitation with transformer- based sensor fusion for autonomous driving. Transactions on Pattern Analysis and Machine Intelligence (T-PAMI), 2023
2023
-
[31]
Karnchanachari, D
N. Karnchanachari, D. Geromichalos, K. S. Tan, N. Li, C. Eriksen, S. Yaghoubi, N. Mehdipour, G. Bernasconi, W. K. Fong, Y . Guo, and H. Caesar. Towards learning-based planning: The nuplan benchmark for real-world autonomous driving. In Proc. IEEE International Conf. on Robotic...
2024
-
[32]
Zheng, R
Y . Zheng, R. Liang, K. Zheng, J. Zheng, L. Mao, J. Li, W. Gu, R. Ai, S. E. Li, X. Zhan, et al. Diffusion- based planning for autonomous driving with flexible guidance. Proc. of the International Conf. on Learning Representations (ICLR), 2025
2025
-
[33]
Carla autonomous driving leaderboard 2.0 scenarios
CARLA. Carla autonomous driving leaderboard 2.0 scenarios. https://leaderboard.carla.org/scenarios, 2022
2022
-
[34]
Leaderboard 1.0 to 2.0 scenario converter
CARLA. Leaderboard 1.0 to 2.0 scenario converter. https://github.com/carla-simulator/leaderboard/ blob/40d507ce67b189e66458a12492acaab706b28e92/scripts/route_bridge.py, 2024
2024
-
[35]
C. team. Carla leaderboard 2.0 metrics. URL : https://leaderboard.carla.org/evaluation_v2_0, 2024
2024
-
[36]
Prakash, A
A. Prakash, A. Behl, E. Ohn-Bar, K. Chitta, and A. Geiger. Exploring data aggregation in policy learn- ing for vision-based urban autonomous driving. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2020
2020
-
[37]
A. Behl, K. Chitta, A. Prakash, E. Ohn-Bar, and A. Geiger. Label efficient visual abstractions for autonomous driving. In Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS), 2020
2020
-
[38]
Islam, P
R. Islam, P. Henderson, M. Gomrokchi, and D. Precup. Reproducibility of benchmarked deep reinforce- ment learning tasks for continuous control. arXiv.org, 1708.04133, 2017
2017 arXiv
-
[39]
Henderson, R
P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger. Deep reinforcement learning that matters. In Proc. of the Conf. on Artificial Intelligence (AAAI), 2018
2018
-
[40]
N. A. Lynnerup, L. Nolling, R. Hasle, and J. Hallam. A survey on reproducibility by evaluating deep reinforcement learning algorithms on real-world robots. InProc. Conf. on Robot Learning (CoRL), 2019
2019
-
[41]
Agarwal, M
R. Agarwal, M. Schwarzer, P. S. Castro, A. C. Courville, and M. G. Bellemare. Deep reinforcement learning at the edge of the statistical precipice. In Advances in Neural Information Processing Systems (NeurIPS), 2021
2021
-
[42]
A. J. Fetterman, E. Kitanidis, J. Albrecht, Z. Polizzi, B. Fogelman, M. Knutins, B. Wróblewski, J. B. Simon, and K. Qiu. Tune as you scale: Hyperparameter optimization for compute efficient training. arXiv.org, 2306.08055, 2023
2023 arXiv
-
[43]
Huang, R
S. Huang, R. F. J. Dossa, A. Raffin, A. Kanervisto, and W. Wang. The 37 implementation details of proximal policy optimization. The ICLR Blog Track 2023, 2023
2023
-
[44]
R. J. Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8:229–256, 1992
1992
-
[45]
V . Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu. Asynchronous methods for deep reinforcement learning. In Proc. of the International Conf. on Machine learning (ICML), 2016
2016
-
[46]
Degris, M
T. Degris, M. White, and R. S. Sutton. Off-policy actor-critic. arXiv.org, 1205.4839, 2012
2012 arXiv
-
[47]
Zhang, J
D. Zhang, J. Liang, K. Guo, S. Lu, Q. Wang, R. Xiong, Z. Miao, and Y . Wang. Carplanner: Consis- tent auto-regressive trajectory planning for large-scale reinforcement learning in autonomous driving. arXiv.org, 2502.19908, 2025
2025 arXiv
-
[48]
V . Koltun. Autonomous driving: The way forward. URL : https://www.youtube.com/watch?v= XmtTjqimW3g, 2020. 11
2020
-
[49]
W. D. Hillis and G. L. S. Jr. Data parallel algorithms. Communications of the ACM, 1986
1986
-
[50]
Kaufmann, L
E. Kaufmann, L. Bauersfeld, A. Loquercio, M. Müller, V . Koltun, and D. Scaramuzza. Champion-level drone racing using deep reinforcement learning. Nature, 2023
2023
-
[51]
Jaeger, K
B. Jaeger, K. Chitta, D. Dauner, K. Renz, and A. Geiger. Common mistakes in benchmarking au- tonomous driving. URL : https://github.com/autonomousvision/carla_garage/blob/leaderboard_2/docs/ common_mistakes_in_benchmarking_ad.md, 2024
2024
-
[52]
Hafner, J
D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap. Mastering diverse control tasks through world models. Nature, 2025
2025
-
[53]
Vaswani, N
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems (NIPS), pages 5998– 6008, 2017
2017
-
[54]
J. Liu, Z. Yang, Z. Huang, W. Li, S. Dang, and H. Li. Simulation performance evaluation of pure pursuit, stanley, lqr, mpc controller for autonomous vehicles. In IEEE international conference on real-time computing and robotics (RCAR). IEEE, 2021
2021
-
[55]
T. Miki, J. Lee, J. Hwangbo, L. Wellhausen, V . Koltun, and M. Hutter. Learning robust perceptive locomotion for quadrupedal robots in the wild. Science Robotics, 2022
2022
-
[56]
T. Lin, K. Sachdev, L. Fan, J. Malik, and Y . Zhu. Sim-to-real reinforcement learning for vision-based dexterous manipulation on humanoids. arXiv.org, 2502.20396, 2025
2025 arXiv
-
[57]
C. J. Watkins and P. Dayan. Q-learning. Machine Learning, 1992
1992
-
[58]
V . Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. A. Ried- miller, A. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Ku- maran, D. Wierstra, S. Legg, and D. Hassabis. Human-level control through d...
2015
-
[59]
Haarnoja, A
T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Proc. of the International Conf. on Machine learning (ICML), 2018
2018
-
[60]
Rajamani
R. Rajamani. Vehicle dynamics and control. Springer Science & Business Media, 2011
2011
-
[61]
Polack, F
P. Polack, F. Altché, B. d’Andréa Novel, and A. de La Fortelle. The kinematic bicycle model: A consistent model for planning feasible trajectories for autonomous vehicles? In Proc. IEEE Intelligent Vehicles Symposium (IV), 2017
2017
-
[62]
Chekroun, T
R. Chekroun, T. Gilles, M. Toromanoff, S. Hornauer, and F. Moutarde. MBAPPE: mcts-built-around prediction for planning explicitly. In Proc. IEEE Intelligent Vehicles Symposium (IV), 2024
2024
-
[63]
B. Yang, H. Su, N. Gkanatsios, T. Ke, A. Jain, J. G. Schneider, and K. Fragkiadaki. Diffusion-es: Gradient-free planning with diffusion for autonomous and instruction-guided driving. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2024
2024
-
[64]
Gulino, J
C. Gulino, J. Fu, W. Luo, G. Tucker, E. Bronstein, Y . Lu, J. Harb, X. Pan, Y . Wang, X. Chen, J. D. Co- Reyes, R. Agarwal, R. Roelofs, Y . Lu, N. Montali, P. Mougin, Z. Yang, B. White, A. Faust, R. McAllister, D. Anguelov, and B. Sapp. Waymax: An accelerated, data-driven simu...
2023
-
[65]
P. Sun, H. Kretzschmar, X. Dotiwalla, A. Chouard, V . Patnaik, P. Tsui, J. Guo, Y . Zhou, Y . Chai, B. Caine, V . Vasudevan, W. Han, J. Ngiam, H. Zhao, A. Timofeev, S. Ettinger, M. Krivokon, A. Gao, A. Joshi, Y . Zhang, J. Shlens, Z. Chen, and D. Anguelov. Scalability in perce...
2020
-
[66]
L. Xiao, J. Liu, X. Ye, W. Yang, and J. Wang. Easychauffeur: A baseline advancing simplicity and efficiency on waymax. arXiv.org, 2408.16375, 2024
2024 arXiv
-
[67]
Charraut, T
V . Charraut, T. Tournaire, W. Doulazmi, and T. Buhet. V-max: Making rl practical for autonomous driving. arXiv.org, 2503.08388, 2025
2025 arXiv
-
[68]
Cusumano-Towner, D
M. Cusumano-Towner, D. Hafner, A. Hertzberg, B. Huval, A. Petrenko, E. Vinitsky, E. Wijmans, T. Kil- lian, S. Bowers, O. Sener, P. Krähenbühl, and V . Koltun. Robust autonomy emerges from self-play. arXiv.org, 2502.03349, 2025. 12
2025 arXiv
-
[69]
S. Ross, G. J. Gordon, and D. Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning. In Conference on Artificial Intelligence and Statistics (AISTATS), 2011
2011
-
[70]
L. Chen, P. Wu, K. Chitta, B. Jaeger, A. Geiger, and H. Li. End-to-end autonomous driving: Challenges and frontiers. Transactions on Pattern Analysis and Machine Intelligence (T-PAMI), 2024
2024
-
[71]
Y . Lu, J. Fu, G. Tucker, X. Pan, E. Bronstein, R. Roelofs, B. Sapp, B. White, A. Faust, S. Whiteson, D. Anguelov, and S. Levine. Imitation is not enough: Robustifying imitation with reinforcement learning for challenging driving scenarios. In Proc. IEEE International Conf. on...
2023
-
[72]
Liang, T
X. Liang, T. Wang, L. Yang, and E. P. Xing. CIRL: controllable imitative reinforcement learning for vision-based self-driving. In Proc. of the European Conf. on Computer Vision (ECCV), 2018
2018
-
[73]
Y . Song, H. Lin, E. Kaufmann, P. Dürr, and D. Scaramuzza. Autonomous overtaking in gran turismo sport using curriculum reinforcement learning. In Proc. IEEE International Conf. on Robotics and Automation (ICRA), 2021
2021
-
[74]
Zhang, A
Z. Zhang, A. Liniger, D. Dai, F. Yu, and L. V . Gool. End-to-end urban driving by imitating a reinforce- ment learning coach. arXiv.org, 2108.08265v3, 2021
2021 arXiv
-
[75]
P. Wu, X. Jia, L. Chen, J. Yan, H. Li, and Y . Qiao. Trajectory-guided control prediction for end-to- end autonomous driving: A simple yet strong baseline. In Advances in Neural Information Processing Systems (NeurIPS), 2022
2022
-
[76]
X. Jia, P. Wu, L. Chen, J. Xie, C. He, J. Yan, and H. Li. Think twice before driving: Towards scalable decoders for end-to-end autonomous driving. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2023
2023
-
[77]
X. Jia, Y . Gao, L. Chen, J. Yan, P. L. Liu, and H. Li. Driveadapter: Breaking the coupling barrier of perception and planning in end-to-end autonomous driving. In Proc. of the IEEE International Conf. on Computer Vision (ICCV), 2023
2023
-
[78]
Ishida, G
S. Ishida, G. Corrado, G. Fedoseev, H. Yeo, L. Russell, J. Shotton, J. F. Henriques, and A. Hu. Langprop: A code optimization framework using large language models applied to driving. InICLR 2024 Workshop on Large Language Model (LLM) Agents, 2024
2024
-
[79]
Huang, R
S. Huang, R. F. J. Dossa, C. Ye, J. Braga, D. Chakraborty, K. Mehta, and J. G. M. Araújo. Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms.Journal of Machine Learning Research (JMLR), 2022
2022
-
[80]
Ha and J
D. Ha and J. Schmidhuber. Recurrent world models facilitate policy evolution. In Advances in Neural Information Processing Systems (NeurIPS), 2018
2018
-
[81]
Hafner, T
D. Hafner, T. P. Lillicrap, J. Ba, and M. Norouzi. Dream to control: Learning behaviors by latent imagination. In Proc. of the International Conf. on Learning Representations (ICLR), 2020
2020
-
[82]
Hafner, T
D. Hafner, T. P. Lillicrap, M. Norouzi, and J. Ba. Mastering atari with discrete world models. In Proc. of the International Conf. on Learning Representations (ICLR), 2021
2021
-
[83]
Kazemkhani, A
S. Kazemkhani, A. Pandya, D. Cornelisse, B. Shacklett, and E. Vinitsky. Gpudrive: Data-driven, multi- agent driving simulation at 1 million FPS. In Proc. of the International Conf. on Learning Representa- tions (ICLR), 2025
2025
-
[84]
Chen and P
D. Chen and P. Krähenbühl. Learning from all vehicles. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2022
2022
-
[85]
J. Weng, M. Lin, S. Huang, B. Liu, D. Makoviichuk, V . Makoviychuk, Z. Liu, Y . Song, T. Luo, Y . Jiang, Z. Xu, and S. Yan. Envpool: A highly parallel reinforcement learning environment execution engine. In Advances in Neural Information Processing Systems (NeurIPS), 2022
2022
-
[86]
Vinyals, I
O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, J. Oh, D. Horgan, M. Kroiss, I. Danihelka, A. Huang, L. Sifre, T. Cai, J. P. Agapiou, M. Jaderberg, A. S. Vezhnevets, R. Leblond, T. Pohlen, V . Dalibard...
2019
-
[87]
Berner, G
C. Berner, G. Brockman, B. Chan, V . Cheung, P. Debiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, R. Józefowicz, S. Gray, C. Olsson, J. Pachocki, M. Petrov, H. P. de Oliveira Pinto, J. Raiman, T. Salimans, J. Schlatter, J. Schneider, S. Sidor, I. Sutskever, J. Ta...
1912 arXiv
-
[88]
Petrenko, Z
A. Petrenko, Z. Huang, T. Kumar, G. S. Sukhatme, and V . Koltun. Sample factory: Egocentric 3d control from pixels at 100000 FPS with asynchronous reinforcement learning. InProc. of the International Conf. on Machine learning (ICML), 2020
2020
-
[89]
Pep 703 – making the global interpreter lock optional in cpython
Python. Pep 703 – making the global interpreter lock optional in cpython. URL : https://peps.python.org/ pep-0703/, 2025
2025
-
[90]
Paszke, S
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Z. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala. Pytorch: An imperative style, high-...
2019
-
[91]
C. Lyle, Z. Zheng, E. Nikishin, B. Á. Pires, R. Pascanu, and W. Dabney. Understanding plasticity in neural networks. In Proc. of the International Conf. on Machine learning (ICML) , Proceedings of Machine Learning Research, 2023
2023
-
[92]
Dohare, J
S. Dohare, J. F. Hernandez-Garcia, Q. Lan, P. Rahman, A. R. Mahmood, and R. S. Sutton. Loss of plasticity in deep continual learning. Nature, 2024
2024
-
[93]
Juliani and J
A. Juliani and J. Ash. A study of plasticity loss in on-policy deep reinforcement learning. In Advances in Neural Information Processing Systems (NeurIPS), 2024
2024
-
[94]
Moalla, A
S. Moalla, A. Miele, D. Pyatko, R. Pascanu, and C. Gulcehre. No representation, no trust: Connecting representation, collapse, and trust issues in PPO. InAdvances in Neural Information Processing Systems (NeurIPS), 2024
2024
-
[95]
P. Chou, D. Maturana, and S. A. Scherer. Improving stochastic policy gradients in continuous control with deep reinforcement learning using the beta distribution. In Proc. of the International Conf. on Machine learning (ICML), 2017
2017
-
[96]
I. G. B. Petrazzini and E. A. Antonelo. Proximal policy optimization with continuous bounded action space via the beta distribution. In IEEE Symposium Series on Computational Intelligence, (SSCI), 2021
2021
-
[97]
Jaeger, K
B. Jaeger, K. Chitta, and A. Geiger. Hidden biases of end-to-end driving models. In Proc. of the IEEE International Conf. on Computer Vision (ICCV), 2023
2023
-
[98]
Brockman, V
G. Brockman, V . Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba. Openai gym. arXiv.org, 1606.01540, 2016
2016 arXiv
-
[99]
Towers, A
M. Towers, A. Kwiatkowski, J. K. Terry, J. U. Balis, G. D. Cola, T. Deleu, M. Goulão, A. Kallinteris, M. Krimmel, A. KG, R. Perez-Vicente, A. Pierré, S. Schulhoff, J. J. Tai, H. Tan, and O. G. Younis. Gymnasium: A standard interface for reinforcement learning environments. arX...
2024 arXiv
-
[100]
Hintjens
P. Hintjens. ZeroMQ: Messaging for Many Applications. O’Reilly Media, 2013
2013
-
[101]
Huang, R
S. Huang, R. F. J. Dossa, A. Raffin, A. Kanervisto, and W. Wang. The 37 implementation details of proximal policy optimization. In ICLR Blog Track, 2022. URL https://iclr-blog-track.github.io/2022/ 03/25/ppo-implementation-details/
2022
-
[102]
Hanselmann, K
N. Hanselmann, K. Renz, K. Chitta, A. Bhattacharyya, and A. Geiger. King: Generating safety-critical driving scenarios for robust imitation via kinematics gradients. In Proc. of the European Conf. on Com- puter Vision (ECCV), 2022
2022
-
[103]
K. Hao, W. Cui, Y . Luo, L. Xie, Y . Bai, J. Yang, S. Yan, Y . Pan, and Z. Yang. Adversarial safety-critical scenario generation using naturalistic human driving priors. IEEE Trans. on Intelligent Transportation Systems (TITS), 2023
2023
-
[104]
Y . Yin, P. Khayatan, Éloi Zablocki, A. Boulch, , and M. Cord. Regents: Real-world safety-critical driving scenario generation made stable. In ECCV 2024 W-CODA Workshop, 2024
2024
-
[105]
D. R. Hipp. Sqlite. URL : http://sqlite.org/, 2000. 14
2000
-
[106]
J. V . den Bossche, K. Jordahl, M. Fleischmann, M. Richards, J. McBride, J. Wasserman, A. G. Badaracco, A. D. Snow, B. Ward, J. Tratner, J. Gerard, M. Perry, cjqf, G. A. Hjelle, M. Taves, E. ter Hoeven, M. Cochran, R. Bell, rraymondgh, M. Bartos, P. Roggemans, L. Culbertson, G...
2024 doi
-
[107]
Zimmerlin, J
J. Zimmerlin, J. Beißwenger, B. Jaeger, A. Geiger, and K. Chitta. Hidden biases of end-to-end driving datasets. arXiv.org, 2412.09602, 2024
2024 arXiv
-
[108]
Hafner, J
D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap. Mastering diverse domains through world models. arXiv.org, 2023
2023
-
[109]
Scheel, L
O. Scheel, L. Bergamini, M. Wolczyk, B. Osi ´nski, and P. Ondruska. Urban driver: Learning to drive from real-world demonstrations using policy gradients. InProc. Conf. on Robot Learning (CoRL), 2021. 15 CaRL: Learning Scalable Planning Policies with Simple Rewards Supplementa...
2021
-
[110]
official
is a 3D open source autonomous driving simulator, based on the Unreal Engine, enabling real- time graphical rendering and simulation. CARLA enables research on the full driving stack from perception to planning and control, and has an extensive set of functionalities extended ...
2020
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.