Pith. sign in

REVIEW 1 major objections 7 minor 5 cited by

CaRL: Learning Scalable Planning Policies with Simple Rewards

T0 review · 1 major / 7 minor · reviewed 2026-08-16 · deepseek-v4-flash

Pith's one-line read Replacing the usual multi-term rewards with a single route-completion term lets a standard RL algorithm scale to 300M samples and reach 64 Driving Score, far above prior RL planners.

desk verdict A genuinely strong RL-for-driving paper whose headline scalability claim is undermined by a missing control; the method and nuPlan results are still worth taking seriously. read the letter →

arxiv 2504.17838 v3 pith:XK2QDROZ submitted 2025-04-24 cs.LG cs.AIcs.RO

classification cs.LGcs.AIcs.RO
keywords reinforcementlearningautonomousdrivingrewarddesignroutecompletionPPOscalingCARLAnuPlanclosed-loopplanning
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

This paper argues that RL planners for autonomous driving have failed to scale because of their reward functions, not their learning algorithms. It shows that under the standard multi-term shaped rewards that add up speed, position, and orientation terms, PPO (Proximal Policy Optimization) collapses when the mini-batch size is increased—getting stuck in local optima such as not driving at all—while a reward made of a single term, route completion, keeps improving as the mini-batch grows. That scaling property unlocks efficient data-parallel training: the authors scale PPO to 300 million samples in CARLA and 500 million in nuPlan on one 8-GPU node, reaching a Driving Score of 64 (route completion multiplied by an infraction penalty) on the longest6 v2 benchmark, 42 points above the prior best RL planner, and the best learning-based closed-loop scores on nuPlan's Val14 benchmark. If this is right, driving planners can be scaled with data the way imitation learners are, without rule-based reward shaping, world models, or hand-tuned trade-offs between reward terms.

What carries the argument

The central object is the single-term reward $r_t = RC_t \left(\prod p_t\right) - T$, where $RC_t$ is the fraction of the route completed during the current simulator step, $p_t \in [0,1]$ are soft penalty factors that multiplicatively shrink the reward while the vehicle violates constraints such as the speed limit, lane-centre distance, time-to-collision, or comfort bounds, and $T$ is a terminal penalty applied at episode end for hard infractions ($T=1$ for collisions and red-light violations, $0$ otherwise). Major infractions terminate the episode, so violating them forfeits all future route completion. The paper motivates three properties of this design: the total reward obtainable is finite (only 100 route-completion points exist), no rule-based planner contributes to any reward term, and the reward's global optimum is the same as the global optimum of the Driving Score metric. The supporting machinery is the scaling stack built around it: Proximal Policy Optimization (PPO) hyperparameters in the Atari style that reduce off-policy gradient steps from 959 to 15, an asynchronous collection loop, DD-PPO (distributed data-parallel PPO) for scaling, and a bird's-eye-view input with a much wider field of view than prior RL planners.

What would settle it

Evaluate the 300M-sample CARLA model on the official CARLA leaderboard 2.0 validation routes (the hand-authored 20-route set on town 13) or on a set of held-out towns never seen in training. If the Driving Score collapses relative to the 64 achieved on longest6 v2—or if a variant trained only on the original hand-crafted scenario definitions matches it—then the auto-generated training distribution, rather than the reward design, is carrying the result. A second, cheaper check replays the mini-batch comparison of Table 3 across seeds: the paper's core claim stands only if the complex reward consistently collapses at mini-batch 1024 while the route-completion reward consistently improves.

Watch

Extended reading notes

Core claim

The paper's central claim is that a driving reward consisting only of route completion, with infractions encoded as episode termination or multiplicative penalties, changes how PPO behaves under scaling: the same algorithm that collapses to 2±2 Driving Score at mini-batch size 1024 with a complex multi-term reward improves to 38±3 with the simple reward, and reaches 64±2 at mini-batch 16384 with 300 million samples. This is presented as evidence that previous RL planners were limited by reward design rather than by data or algorithm: the complex reward's handcrafted terms introduce local minima and rule-based upper bounds on performance, whereas the route-completion reward has a finite reward budget (only 100 route-completion points exist), contains no rule-based planner, and shares its global optimum with the evaluation metric. On nuPlan the identical recipe, adapted only by adding a survival bonus for the simulator's fixed-length episodes, achieves 91.3 closed-loop score in non-reactive and 90.6 in reactive traffic on the Val14 benchmark, making it the best learning-based method while running an order of magnitude faster at inference than prior work.

Load-bearing premise

The load-bearing premise is that the training routes and scenarios generated automatically by rejection sampling and heuristics, rather than by human labelling, are diverse and representative enough that a policy trained on them transfers to the fixed evaluation routes of longest6 v2; if those generated scenarios are easier or biased relative to the test routes, the reported Driving Scores would overstate the policy's real planning ability.

Editorial extensions

If this is right

  • RL planners for driving can be scaled with data the way imitation learners are: under the simple reward, raising the sample count from 10 million to 300 million with a larger mini-batch increases Driving Score by 33 points, the largest single improvement reported in this line of work.
  • Rule-based planners can be dropped from the reward pipeline entirely: the resulting policy outperforms every prior RL planner on CARLA and every learning-based planner on nuPlan without any scenario-specific rules.
  • Large mini-batch sizes become an asset rather than a liability for on-policy RL in driving, so distributed data-parallel collection (DD-PPO) is the right scaling regime for this task.
  • The same reward and hyperparameters transfer across simulators with only a survival-bonus adaptation, suggesting the design exploits a general property of goal-directed driving rather than CARLA-specific quirks.
  • Because the reward's global optimum equals the evaluation metric's, further gains should come from more data and model capacity rather than reward re-tuning; the remaining gap to the rule-based PDM-Lite is attributed to consistency in solving safety-critical scenarios, not to reward misalignment.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The property that likely carries the scaling result is the finiteness of the reward budget: with only 100 route-completion points available, no stationary behaviour can harvest reward forever, which is exactly the loophole the paper documents in the complex reward (waiting at green lights); any long-horizon task with a natural 'fraction of goal achieved' signal—navigation, manipulation, search—may
  • The paper's explanation for the complex-reward collapse—that larger mini-batches smooth the gradient and push early optimization into the broadest local optimum (not driving)—is testable directly: measuring per-term gradient correlations or loss-landscape curvature as a function of mini-batch size would confirm or refute it without any new hardware.
  • Because training and evaluation share the same towns (level 4), the natural next claim to test is generalization: the reward and inputs carry no town-specific rules, so the recipe is a plausible starting point for zero-shot transfer to unseen towns, which the paper explicitly leaves open.
  • The nuPlan survival bonus is a workaround for episodes of fixed length; in open-ended operation with no defined route end, route completion alone would stop giving signal once the goal is reached, so any deployment beyond benchmarks would need a terminal-state design or an explicit 'keep driving safely' term.
Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

1 major / 7 minor

Summary. The paper proposes CaRL, a privileged RL planning method for CARLA and nuPlan trained with PPO and a sparse reward that rewards route completion, terminates episodes on hard infractions, and applies multiplicative soft penalties for constraint violations. The authors report that this reward, unlike a shaped Roach-style reward, allows PPO to benefit from larger mini-batch sizes, enabling scaling to 300M samples in CARLA and 500M/1B samples in nuPlan on a single 8-GPU node. They report 64 DS on the longest6 v2 benchmark, outperforming prior RL planners by a large margin, and 91.3/90.6 CLS on nuPlan Val14, outperforming prior learning-based planners. The paper includes ablations with multiple seeds, a released codebase, and reimplementations of baselines.

Significance. If the central scalability claim holds, this is a noteworthy result for RL-based driving: it challenges the prevailing shaped-reward paradigm, demonstrates that a metric-aligned sparse reward can be optimized at scale with PPO, and provides a reproducible recipe for CARLA and nuPlan. The engineering contributions, including the RL-optimized leaderboard code, AC-PPO, and large-scale training, are also valuable. The paper's reproducibility practices are strong: code is released, experiments average multiple seeds, baselines are reimplemented, and evaluation uses the official leaderboard code. However, the key comparison used to attribute the scalability advantage to reward design is confounded, so the significance of the conceptual claim is currently uncertain pending a matched-update experiment.

major comments (1)
  1. [Section 3, Table 3; Section 3.1] Section 3.1 states that the mini-batch 1024 runs were obtained by 'performing 4x more simulator steps per PPO iteration with 4x fewer PPO iterations' while holding total samples fixed. This means the comparison in Table 3 changes three variables simultaneously: mini-batch size (256 to 1024), rollout length per PPO iteration (128 to 512), and total number of optimizer updates (reduced by a factor of 4). Longer rollouts directly improve return estimation for a sparse reward, and fewer updates can explain the collapse of the Roach reward (34±7 to 2±2) as undertraining rather than a reward-induced local minimum. The text's attribution of the collapse to 'larger mini-batch sizes smooth the optimization' (Section 3) is therefore not supported by the evidence presented. Please run an ablation that varies mini-batch size while holding rollout length and total gradient updates fixed, e.g., by increasing the number of parallel environments instead of the rollout length and keeping the number of PPO iterations constant; report learning curves and total update counts. If the complex reward recovers with a matched update budget, the abstract's claim that complex rewards 'limit scalability' should be revised to a compute-budget effect rather than a reward-design effect.
minor comments (7)
  1. [Section 3, Eq. (1), principle (3)] The statement that 'the global optimum of the reward is the same as the global optimum of the metric' is an overstatement: the reward includes soft penalties for speeding, TTC, and comfort that are not part of the CARLA DS metric, so a policy that maximizes DS but violates these soft constraints would receive suboptimal reward. Please soften this to 'closely aligned' and state the residual differences.
  2. [Section 3.1, Table 3] Please state in the table caption or text the rollout length and total number of gradient updates used in each column; currently the reader cannot disentangle mini-batch size from these other changes.
  3. [Abstract and Introduction] The phrase 'single intuitive reward term' is somewhat misleading because Eq. (1) includes several multiplicative soft penalties and terminal penalties; consider describing it as a single primary reward term with constraint penalties.
  4. [Table 4 and Section 3.1] The 300M model is averaged over three seeds while other experiments use five; please justify the reduced number of seeds for the main result.
  5. [Section 4.2, Table 6] The claim of being the best learning-based method on nuPlan should be qualified: PLUTO (92.6 NR) is excluded as a hybrid IL+rule method, and the main-text comparison does not include it in the table. Please make the exclusion criterion explicit and consider reporting PLUTO in Table 6.
  6. [Appendix D.2] The automatically generated training routes and scenarios are acknowledged to be noisy and unbalanced; adding basic statistics, such as number of routes, scenario type distribution, and rejection rates, would help assess the generalization evidence.
  7. [Appendix E.1.2] Because Think2Drive's code is not public and key hyperparameters were inferred, the reported 7 DS result may be sensitive to these choices; please provide a short sensitivity analysis or release the reproduction to strengthen the comparison.

Circularity Check

0 steps flagged · score 2.0 of 10

No significant circularity: reward–metric alignment is a stated design choice, the scaling claim is empirical, and the nuPlan-tuned variant is disclosed rather than presented as a prediction.

full rationale

The central derivation—that a route-completion reward with terminal/soft penalties allows PPO to scale to large mini-batches and 300M samples—is an empirical result, not a consequence of definition. Section 3 explicitly states design principle (3) that the reward's global optimum matches the metric's optimum; this is a deliberate objective alignment (standard reward design) rather than a hidden equivalence, and the paper confirms the reward 'closely aligns with the DS metric' while adding extra penalties. The reported CARLA DS is not a fitted parameter: no aspect of the reward or architecture is tuned to longest6 v2 scores beyond the stated reward form, and comparisons to Roach/Think2Drive/PDM-Lite are independent reproductions. The nuPlan 'CaRL (nuPlan tuned)' variant in Appendix E.2 does take the CLS metric's own 5:5:4:2 weights and add TTC/comfort/speed terms to route completion, so its improved CLS is partly by construction; however, the paper labels it 'nuPlan tuned', discloses the exact weight copying, warns that the reward is not recommended generally, and does not use it for the headline 91.3/90.6 claims. Self-citations ([17] for reward–metric alignment, [51] for benchmarking practice) are not load-bearing: the alignment principle is domain-general and the benchmark results stand on the paper's own experiments. The Table 3 mini-batch comparison confounds mini-batch size with number of PPO updates (4x fewer iterations at 1024), but this is a validity/confound issue rather than circularity, and it is not a case where a fitted input is renamed as a prediction. Overall, no derivation step reduces to its own input by construction.

Assumptions & free parameters 8 free parameters · 5 assumptions · 0 invented entities

The paper's central claims rest on a small number of hand-set reward parameters, all disclosed, and on domain assumptions about simulator fidelity and the sufficiency of the observation. No new physical entities are postulated. The reward parameter values are not fitted to the test benchmark, but they are free choices; different choices could change the results.

free parameters (8)
  • survival_ratio_s = 0.6 (nuPlan)
    Weight for combining route-completion reward with a constant survival bonus in nuPlan (Eq. 8). Chosen by the authors; not fitted to maximize the benchmark.
  • terminal_penalty_T = 1.0 for collision/red light, 0.0 otherwise
    Sets the ordering of terminal penalties; chosen to avoid discouraging the agent from driving early in training (Section F.1).
  • soft_penalty_factors = pt = 0 (outside lanes), 0.5 (TTC), 1 - 0.5*#/6 (comfort)
    Multiplicative penalties for soft infractions (Section F.2).
  • comfort_thresholds = Table 16 (e.g., longitudinal accel (-20,10) m/s^2 in CARLA)
    Hand-set bounds for comfort infractions; paper acknowledges lack of human reference and sets wide bounds.
  • route_deviation_threshold = 30 m
    Terminal threshold for route deviation in CARLA (Section F.1).
  • blocked_threshold = 90 s at <0.1 m/s
    Terminal penalty for blocked episodes (Section F.1).
  • action_scaling_nuPlan = [-3.2, 2.4] m/s^2, [-0.84, 0.84] rad
    Scaling of action outputs to the nuPlan bicycle model; derived from applying default controller to human trajectories (Section B.6).
  • train_frequency = 10 Hz
    Simulator frequency during training; lower 20 Hz eval with action repeat 2 (Section D.1).
assumptions (5)
  • domain assumption The CARLA and nuPlan simulators provide sufficiently faithful models of urban driving such that policies optimized in them transfer to the evaluation protocol.
    The entire evaluation rests on simulator fidelity; if the simulators are unrealistic, the scores are not meaningful for real driving. Stated in Section 2.1 and used throughout.
  • domain assumption The BEV semantic observation, with the modifications in Section B.5, contains sufficient information to drive safely.
    The policy is a function of this input; if critical information is missing (e.g., FOV too small), performance is limited.
  • domain assumption PPO with the Atari hyperparameters is a sufficiently capable optimizer for the simple reward at scale.
    The paper relies on PPO's ability to optimize a sparse reward with Monte-Carlo returns; this is an empirical assumption validated by results but not proven.
  • domain assumption The global optimum of the proposed reward (Eq. 1) aligns with the global optimum of the driving metric (DS/CLS).
    This design principle is stated in Section 3; if misaligned, optimizing the reward would not yield good metric scores.
  • domain assumption The automatically generated training routes and scenarios are representative of the evaluation distribution.
    Section D.2 describes rejection sampling; if the generation is biased, the policy may overfit to easy routes.

how reviews work

0 comments
Cite this review

Pith. "Pith review of CaRL: Learning Scalable Planning Policies with Simple Rewards." pith.science (2026). https://pith.science/paper/XK2QDROZ

@misc{pith2026250417838,
  author       = {Pith},
  title        = {Pith review of: CaRL: Learning Scalable Planning Policies with Simple Rewards},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/XK2QDROZ}},
  note         = {Machine review of arXiv:2504.17838}
}
read the original abstract

We investigate reinforcement learning (RL) for privileged planning in autonomous driving. State-of-the-art approaches for this task are rule-based, but these methods do not scale to the long tail. RL, on the other hand, is scalable and does not suffer from compounding errors like imitation learning. Contemporary RL approaches for driving use complex shaped rewards that sum multiple individual rewards, \eg~progress, position, or orientation rewards. We show that PPO fails to optimize a popular version of these rewards when the mini-batch size is increased, which limits the scalability of these approaches. Instead, we propose a new reward design based primarily on optimizing a single intuitive reward term: route completion. Infractions are penalized by terminating the episode or multiplicatively reducing route completion. We find that PPO scales well with higher mini-batch sizes when trained with our simple reward, even improving performance. Training with large mini-batch sizes enables efficient scaling via distributed data parallelism. We scale PPO to 300M samples in CARLA and 500M samples in nuPlan with a single 8-GPU node. The resulting model achieves 64 DS on the CARLA longest6 v2 benchmark, outperforming other RL methods with more complex rewards by a large margin. Requiring only minimal adaptations from its use in CARLA, the same method is the best learning-based approach on nuPlan. It scores 91.3 in non-reactive and 90.6 in reactive traffic on the Val14 benchmark while being an order of magnitude faster than prior work.

Figures

Figures reproduced from arXiv: 2504.17838 by the authors.

Figure 1
Figure 1. Simple rewards scale with mini-batch size. Typical rewards in driving consist of com￾plex rewards that trade off many individual components. This limits scalability as PPO gets stuck in local minima with larger mini-batch sizes. We propose a simple alternative based on maximizing route completion that scales well with mini-batch size. In settings where other dynamic actors are present, which require more complex beh… view at source ↗
Figure 2
Figure 2. Beta distribution examples. B.5 Input Similar to prior work [14, 27], we use a BEV semantic segmentation image as input. The channels consist of a channel for the road, an A⋆ route mask, lane markings, vehicles, pedestrians, traffic lights, stop signs, speed signs, static objects, and shoulder lanes. Unlike prior work [14, 27], we do not use additional channels for the states of pedestrians and vehicles from past ti… view at source ↗
Figure 3
Figure 3. A rendering of our input. The distribution shows the model action distribution predic￾tions. The yellow vertical line denotes the mean of the distribution. Other cars are rendered in blue. The brightness of blue encodes their speed. A constant velocity forecast is rendered as a line in front of other vehicles. The ego car is depicted in white. Conditioning is in light grey and only rendered inside intersections. Dar… view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Roach lane changes at predefined locations leading to rear collisions. White is the ego car. Blue other cars. Lighter shades of blue represent past time steps. Gray is the precomputed route from the A⋆ planner, and decides which lane to drive on [PITH_FULL_IMAGE:figur…
Figure 5
Figure 5. Figure 5: Roach learns to wait at a green light early during training. Note the 0 m/s velocity at the top of the image. drives at 4 m/s, slows down to 3 m/s when it gets close to the traffic light. Once it passes the traffic light, the model starts accelerating again here to 5 m…
Figure 6
Figure 6. Figure 6: Roach slows down at green lights. Top to bottom are 3 different time steps. Note how the velocity initially slows down from 4 m/s to 3 m/s. Once the agent passed the green light, the model started accelerating again to 5 m/s. 36 [PITH_FULL_IMAGE:figures/full_fig_p036_6.png]
Figure 7
Figure 7. Figure 7: Another car runs a red light and rams the CaRL from behind [PITH_FULL_IMAGE:figures/full_fig_p037_7.png]
Figure 8
Figure 8. Figure 8: CaRL misses a highway exit. (a) Pedestrian collision. (b) Vehicle collision. (c) Ego stuck in traffic [PITH_FULL_IMAGE:figures/full_fig_p037_8.png]
Figure 9
Figure 9. Figure 9: Failures of CaRL in nuPlan. (a) The ego vehicle collides with non-reactive pedestrians and veers off-road in an attempt at collision avoidance. (b) During an unprotected turn, the ego vehicle has a rear-side collision with a non-reactive vehicle. (c) Non-ego vehicles r…
Figure 10
Figure 10. Figure 10: Training examples with the updated PlanT inputs. The image on the left shows the ego [PITH_FULL_IMAGE:figures/full_fig_p038_10.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 5 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. On Data Thinning for Model Validation in Small Area Estimation

    stat.ME 2026-04 unverdicted novelty 7.0 of 10

    Thinned-data MSE for small-area models is unbiased for a risk that systematically differs from full-data risk; under Fay-Herriot the gap is closed-form in the model's shrinkage, and the thinning fraction faces a sharp...

  2. Threat-guided Policy-aware Scene Perturbation for Safe Autonomous Driving with Online Reinforcement Learning

    cs.AI 2026-08 conditional novelty 6.0 of 10

    TPSP trains a scene-perturbation network using threat differences between original and perturbed rollouts to create safety-critical experiences for an online RL driving policy.

  3. Zero-Human Demonstration End-to-end Autonomous Driving with Trajectory Scorer

    cs.RO 2025-10 conditional novelty 6.0 of 10

    A reward-only offline RL method for trajectory planning in end-to-end autonomous driving achieves state-of-the-art on Navhard and competitive closed-loop HUGSIM performance without imitation learning.

  4. CLEAR: Closed-Loop Reinforcement Learning at Scale for End-to-End Autonomous Driving

    cs.RO 2026-07 conditional novelty 5.5 of 10

    Residual waypoint RL around a frozen VLA prior, scaled via heterogeneous CARLA/H100 infrastructure, raises closed-loop driving score and success rate on longest6 v2 and Bench2Drive.

  5. End-to-End Crop Row Navigation via LiDAR-Based Deep Reinforcement Learning

    cs.RO 2025-09 conditional novelty 5.0 of 10

    Raw 3D LiDAR, compressed into flattened voxel maps, trains a reinforcement learning policy that reliably follows straight crop rows in simulation and degrades on curvier rows.

Reference graph

Works this paper leans on

110 extracted references · 58 canonical work pages · cited by 5 Pith papers

  1. [1]

    Treiber, A

    M. Treiber, A. Hennecke, and D. Helbing. Congested traffic states in empirical observations and micro- scopic simulations. Physical review E, 2000

  2. [2]

    Thrun, M

    S. Thrun, M. Montemerlo, H. Dahlkamp, D. Stavens, A. Aron, J. Diebel, P. Fong, J. Gale, M. Halpenny, G. Hoffmann, K. Lau, C. M. Oakley, M. Palatucci, V . R. Pratt, P. Stang, S. Strohband, C. Dupont, L. Jendrossek, C. Koelen, C. Markey, C. Rummel, J. van Niekerk, E. Jensen, P. Alessandrini, G. R. Bradski, B. Davies, S. Ettinger, A. Kaehler, A. V . Nefian, ...

  3. [3]

    B. Jaeger. Expert drivers for autonomous driving. Master’s thesis, University of Tübingen, 2021

  4. [4]

    Dauner, M

    D. Dauner, M. Hallgarten, A. Geiger, and K. Chitta. Parting with misconceptions about learning-based vehicle motion planning. In Proc. Conf. on Robot Learning (CoRL), 2023

  5. [5]

    C. Sima, K. Renz, K. Chitta, L. Chen, H. Zhang, C. Xie, J. Beißwenger, P. Luo, A. Geiger, and H. Li. Drivelm: Driving with graph visual question answering. In Proc. of the European Conf. on Computer Vision (ECCV), 2024

  6. [6]

    K. Renz, K. Chitta, O.-B. Mercea, A. S. Koepke, Z. Akata, and A. Geiger. Plant: Explainable planning transformers via object-level representations. In Proc. Conf. on Robot Learning (CoRL), 2022. 9

  7. [7]

    Hallgarten, M

    M. Hallgarten, M. Stoll, and A. Zell. From prediction to planning with goal conditioned lane graph traversals. In Proc. IEEE Conf. on Intelligent Transportation Systems (ITSC), 2023

  8. [8]

    Cheng, Y

    J. Cheng, Y . Chen, X. Mei, B. Yang, B. Li, and M. Liu. Rethinking imitation-based planners for au- tonomous driving. In Proc. IEEE International Conf. on Robotics and Automation (ICRA), 2024

Show all 110 references
  1. [9]

    Huang, H

    Z. Huang, H. Liu, J. Wu, and C. Lv. Differentiable integrated motion prediction and planning with learnable cost function for autonomous driving. IEEE Trans. Neural Networks Learn. Syst., 2024

  2. [10]

    Q. Sun, H. Wang, J. Zhan, F. Nie, X. Wen, L. Xu, K. Zhan, P. Jia, X. Lang, and H. Zhao. Generalizing motion planners with mixture of experts for autonomous driving. arXiv.org, 2410.15774, 2024

  3. [11]

    J. Guo, M. Feng, P. Zhu, C. Li, and J. Pu. Rethinking closed-loop planning framework for imitation- based model integrating prediction and planning. arXiv.org, 2407.05376, 2024

  4. [12]

    Cheng, Y

    J. Cheng, Y . Chen, and Q. Chen. PLUTO: pushing the limit of imitation learning-based planning for autonomous driving. arXiv.org, 2404.14327, 2024

  5. [13]

    X. Chen, J. Yan, W. Liao, T. He, and P. Peng. Int2planner: An intention-based multi-modal motion planner for integrated prediction and planning. arXiv.org, 2501.12799, 2025

  6. [14]

    Zhang, A

    Z. Zhang, A. Liniger, D. Dai, F. Yu, and L. Van Gool. End-to-end urban driving by imitating a reinforce- ment learning coach. In Proc. of the IEEE International Conf. on Computer Vision (ICCV), 2021

  7. [15]

    Zhang, R

    C. Zhang, R. Guo, W. Zeng, Y . Xiong, B. Dai, R. Hu, M. Ren, and R. Urtasun. Rethinking closed-loop training for autonomous driving. In Proc. of the European Conf. on Computer Vision (ECCV), 2022

  8. [16]

    R. S. Sutton and A. G. Barto. Reinforcement learning: An introduction. MIT press, 2018

  9. [17]

    Jaeger and A

    B. Jaeger and A. Geiger. An invitation to deep reinforcement learning. Foundations and Trends in Optimization, 2024

  10. [18]

    W. B. Knox, A. Allievi, H. Banzhaf, F. Schmitt, and P. Stone. Reward (mis)design for autonomous driving. Artificial Intelligence (AI), 2023

  11. [19]

    Kendall, J

    A. Kendall, J. Hawke, D. Janz, P. Mazur, D. Reda, J. Allen, V . Lam, A. Bewley, and A. Shah. Learning to drive in a day. In Proc. IEEE International Conf. on Robotics and Automation (ICRA), 2019

  12. [20]

    Wijmans, A

    E. Wijmans, A. Kadian, A. Morcos, S. Lee, I. Essa, D. Parikh, M. Savva, and D. Batra. DD-PPO: learning near-perfect pointgoal navigators from 2.5 billion frames. In Proc. of the International Conf. on Learning Representations (ICLR), 2020

  13. [21]

    Fuchs, Y

    F. Fuchs, Y . Song, E. Kaufmann, D. Scaramuzza, and P. Dürr. Super-human performance in gran turismo sport using deep reinforcement learning. IEEE Robotics and Automation Letters (RA-L), 2021

  14. [22]

    K. Zeng, Z. Zhang, K. Ehsani, R. Hendrix, J. Salvador, A. Herrasti, R. B. Girshick, A. Kembhavi, and L. Weihs. Poliformer: Scaling on-policy RL with transformers results in masterful navigators. In Proc. Conf. on Robot Learning (CoRL), volume 2406.20083, 2024

  15. [23]

    P. R. Wurman, S. Barrett, K. Kawamoto, J. MacGlashan, K. Subramanian, T. J. Walsh, R. Capobianco, A. Devlic, F. Eckert, F. Fuchs, L. Gilpin, P. Khandelwal, V . Kompella, H. Lin, P. MacAlpine, D. Oller, T. Seno, C. Sherstan, M. D. Thomure, H. Aghabozorgi, L. Barrett, R. Douglas...

  16. [24]

    Schulman, F

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algo- rithms. arXiv.org, 1707.06347, 2017

  17. [25]

    Toromanoff, E

    M. Toromanoff, E. Wirbel, and F. Moutarde. End-to-end model-free reinforcement learning for urban driving using implicit affordances. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2020

  18. [26]

    Chekroun, M

    R. Chekroun, M. Toromanoff, S. Hornauer, and F. Moutarde. GRI: general reinforced imitation and its application to vision-based autonomous driving. Robotics, 2023

  19. [27]

    Q. Li, X. Jia, S. Wang, and J. Yan. Think2drive: Efficient reinforcement learning by thinking with latent world model for autonomous driving (in CARLA-V2). In Proc. of the European Conf. on Computer Vision (ECCV), 2024. 10

  20. [28]

    Dosovitskiy, G

    A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, and V . Koltun. CARLA: An open urban driving simu- lator. In Proc. Conf. on Robot Learning (CoRL), 2017

  21. [29]

    C. team. Carla autonomous driving leaderboard 2.0. https://leaderboard.carla.org/, 2022

  22. [30]

    Chitta, A

    K. Chitta, A. Prakash, B. Jaeger, Z. Yu, K. Renz, and A. Geiger. Transfuser: Imitation with transformer- based sensor fusion for autonomous driving. Transactions on Pattern Analysis and Machine Intelligence (T-PAMI), 2023

  23. [31]

    Karnchanachari, D

    N. Karnchanachari, D. Geromichalos, K. S. Tan, N. Li, C. Eriksen, S. Yaghoubi, N. Mehdipour, G. Bernasconi, W. K. Fong, Y . Guo, and H. Caesar. Towards learning-based planning: The nuplan benchmark for real-world autonomous driving. In Proc. IEEE International Conf. on Robotic...

  24. [32]

    Zheng, R

    Y . Zheng, R. Liang, K. Zheng, J. Zheng, L. Mao, J. Li, W. Gu, R. Ai, S. E. Li, X. Zhan, et al. Diffusion- based planning for autonomous driving with flexible guidance. Proc. of the International Conf. on Learning Representations (ICLR), 2025

  25. [33]

    Carla autonomous driving leaderboard 2.0 scenarios

    CARLA. Carla autonomous driving leaderboard 2.0 scenarios. https://leaderboard.carla.org/scenarios, 2022

  26. [34]

    Leaderboard 1.0 to 2.0 scenario converter

    CARLA. Leaderboard 1.0 to 2.0 scenario converter. https://github.com/carla-simulator/leaderboard/ blob/40d507ce67b189e66458a12492acaab706b28e92/scripts/route_bridge.py, 2024

  27. [35]

    C. team. Carla leaderboard 2.0 metrics. URL : https://leaderboard.carla.org/evaluation_v2_0, 2024

  28. [36]

    Prakash, A

    A. Prakash, A. Behl, E. Ohn-Bar, K. Chitta, and A. Geiger. Exploring data aggregation in policy learn- ing for vision-based urban autonomous driving. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2020

  29. [37]

    A. Behl, K. Chitta, A. Prakash, E. Ohn-Bar, and A. Geiger. Label efficient visual abstractions for autonomous driving. In Proc. IEEE International Conf. on Intelligent Robots and Systems (IROS), 2020

  30. [38]

    Islam, P

    R. Islam, P. Henderson, M. Gomrokchi, and D. Precup. Reproducibility of benchmarked deep reinforce- ment learning tasks for continuous control. arXiv.org, 1708.04133, 2017

  31. [39]

    Henderson, R

    P. Henderson, R. Islam, P. Bachman, J. Pineau, D. Precup, and D. Meger. Deep reinforcement learning that matters. In Proc. of the Conf. on Artificial Intelligence (AAAI), 2018

  32. [40]

    N. A. Lynnerup, L. Nolling, R. Hasle, and J. Hallam. A survey on reproducibility by evaluating deep reinforcement learning algorithms on real-world robots. InProc. Conf. on Robot Learning (CoRL), 2019

  33. [41]

    Agarwal, M

    R. Agarwal, M. Schwarzer, P. S. Castro, A. C. Courville, and M. G. Bellemare. Deep reinforcement learning at the edge of the statistical precipice. In Advances in Neural Information Processing Systems (NeurIPS), 2021

  34. [42]

    A. J. Fetterman, E. Kitanidis, J. Albrecht, Z. Polizzi, B. Fogelman, M. Knutins, B. Wróblewski, J. B. Simon, and K. Qiu. Tune as you scale: Hyperparameter optimization for compute efficient training. arXiv.org, 2306.08055, 2023

  35. [43]

    Huang, R

    S. Huang, R. F. J. Dossa, A. Raffin, A. Kanervisto, and W. Wang. The 37 implementation details of proximal policy optimization. The ICLR Blog Track 2023, 2023

  36. [44]

    R. J. Williams. Simple statistical gradient-following algorithms for connectionist reinforcement learning. Machine Learning, 8:229–256, 1992

  37. [45]

    V . Mnih, A. P. Badia, M. Mirza, A. Graves, T. P. Lillicrap, T. Harley, D. Silver, and K. Kavukcuoglu. Asynchronous methods for deep reinforcement learning. In Proc. of the International Conf. on Machine learning (ICML), 2016

  38. [46]

    Degris, M

    T. Degris, M. White, and R. S. Sutton. Off-policy actor-critic. arXiv.org, 1205.4839, 2012

  39. [47]

    Zhang, J

    D. Zhang, J. Liang, K. Guo, S. Lu, Q. Wang, R. Xiong, Z. Miao, and Y . Wang. Carplanner: Consis- tent auto-regressive trajectory planning for large-scale reinforcement learning in autonomous driving. arXiv.org, 2502.19908, 2025

  40. [48]

    V . Koltun. Autonomous driving: The way forward. URL : https://www.youtube.com/watch?v= XmtTjqimW3g, 2020. 11

  41. [49]

    W. D. Hillis and G. L. S. Jr. Data parallel algorithms. Communications of the ACM, 1986

  42. [50]

    Kaufmann, L

    E. Kaufmann, L. Bauersfeld, A. Loquercio, M. Müller, V . Koltun, and D. Scaramuzza. Champion-level drone racing using deep reinforcement learning. Nature, 2023

  43. [51]

    Jaeger, K

    B. Jaeger, K. Chitta, D. Dauner, K. Renz, and A. Geiger. Common mistakes in benchmarking au- tonomous driving. URL : https://github.com/autonomousvision/carla_garage/blob/leaderboard_2/docs/ common_mistakes_in_benchmarking_ad.md, 2024

  44. [52]

    Hafner, J

    D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap. Mastering diverse control tasks through world models. Nature, 2025

  45. [53]

    Vaswani, N

    A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems (NIPS), pages 5998– 6008, 2017

  46. [54]

    J. Liu, Z. Yang, Z. Huang, W. Li, S. Dang, and H. Li. Simulation performance evaluation of pure pursuit, stanley, lqr, mpc controller for autonomous vehicles. In IEEE international conference on real-time computing and robotics (RCAR). IEEE, 2021

  47. [55]

    T. Miki, J. Lee, J. Hwangbo, L. Wellhausen, V . Koltun, and M. Hutter. Learning robust perceptive locomotion for quadrupedal robots in the wild. Science Robotics, 2022

  48. [56]

    T. Lin, K. Sachdev, L. Fan, J. Malik, and Y . Zhu. Sim-to-real reinforcement learning for vision-based dexterous manipulation on humanoids. arXiv.org, 2502.20396, 2025

  49. [57]

    C. J. Watkins and P. Dayan. Q-learning. Machine Learning, 1992

  50. [58]

    V . Mnih, K. Kavukcuoglu, D. Silver, A. A. Rusu, J. Veness, M. G. Bellemare, A. Graves, M. A. Ried- miller, A. Fidjeland, G. Ostrovski, S. Petersen, C. Beattie, A. Sadik, I. Antonoglou, H. King, D. Ku- maran, D. Wierstra, S. Legg, and D. Hassabis. Human-level control through d...

  51. [59]

    Haarnoja, A

    T. Haarnoja, A. Zhou, P. Abbeel, and S. Levine. Soft actor-critic: Off-policy maximum entropy deep reinforcement learning with a stochastic actor. In Proc. of the International Conf. on Machine learning (ICML), 2018

  52. [60]

    Rajamani

    R. Rajamani. Vehicle dynamics and control. Springer Science & Business Media, 2011

  53. [61]

    Polack, F

    P. Polack, F. Altché, B. d’Andréa Novel, and A. de La Fortelle. The kinematic bicycle model: A consistent model for planning feasible trajectories for autonomous vehicles? In Proc. IEEE Intelligent Vehicles Symposium (IV), 2017

  54. [62]

    Chekroun, T

    R. Chekroun, T. Gilles, M. Toromanoff, S. Hornauer, and F. Moutarde. MBAPPE: mcts-built-around prediction for planning explicitly. In Proc. IEEE Intelligent Vehicles Symposium (IV), 2024

  55. [63]

    B. Yang, H. Su, N. Gkanatsios, T. Ke, A. Jain, J. G. Schneider, and K. Fragkiadaki. Diffusion-es: Gradient-free planning with diffusion for autonomous and instruction-guided driving. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2024

  56. [64]

    Gulino, J

    C. Gulino, J. Fu, W. Luo, G. Tucker, E. Bronstein, Y . Lu, J. Harb, X. Pan, Y . Wang, X. Chen, J. D. Co- Reyes, R. Agarwal, R. Roelofs, Y . Lu, N. Montali, P. Mougin, Z. Yang, B. White, A. Faust, R. McAllister, D. Anguelov, and B. Sapp. Waymax: An accelerated, data-driven simu...

  57. [65]

    P. Sun, H. Kretzschmar, X. Dotiwalla, A. Chouard, V . Patnaik, P. Tsui, J. Guo, Y . Zhou, Y . Chai, B. Caine, V . Vasudevan, W. Han, J. Ngiam, H. Zhao, A. Timofeev, S. Ettinger, M. Krivokon, A. Gao, A. Joshi, Y . Zhang, J. Shlens, Z. Chen, and D. Anguelov. Scalability in perce...

  58. [66]

    L. Xiao, J. Liu, X. Ye, W. Yang, and J. Wang. Easychauffeur: A baseline advancing simplicity and efficiency on waymax. arXiv.org, 2408.16375, 2024

  59. [67]

    Charraut, T

    V . Charraut, T. Tournaire, W. Doulazmi, and T. Buhet. V-max: Making rl practical for autonomous driving. arXiv.org, 2503.08388, 2025

  60. [68]

    Cusumano-Towner, D

    M. Cusumano-Towner, D. Hafner, A. Hertzberg, B. Huval, A. Petrenko, E. Vinitsky, E. Wijmans, T. Kil- lian, S. Bowers, O. Sener, P. Krähenbühl, and V . Koltun. Robust autonomy emerges from self-play. arXiv.org, 2502.03349, 2025. 12

  61. [69]

    S. Ross, G. J. Gordon, and D. Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning. In Conference on Artificial Intelligence and Statistics (AISTATS), 2011

  62. [70]

    L. Chen, P. Wu, K. Chitta, B. Jaeger, A. Geiger, and H. Li. End-to-end autonomous driving: Challenges and frontiers. Transactions on Pattern Analysis and Machine Intelligence (T-PAMI), 2024

  63. [71]

    Y . Lu, J. Fu, G. Tucker, X. Pan, E. Bronstein, R. Roelofs, B. Sapp, B. White, A. Faust, S. Whiteson, D. Anguelov, and S. Levine. Imitation is not enough: Robustifying imitation with reinforcement learning for challenging driving scenarios. In Proc. IEEE International Conf. on...

  64. [72]

    Liang, T

    X. Liang, T. Wang, L. Yang, and E. P. Xing. CIRL: controllable imitative reinforcement learning for vision-based self-driving. In Proc. of the European Conf. on Computer Vision (ECCV), 2018

  65. [73]

    Y . Song, H. Lin, E. Kaufmann, P. Dürr, and D. Scaramuzza. Autonomous overtaking in gran turismo sport using curriculum reinforcement learning. In Proc. IEEE International Conf. on Robotics and Automation (ICRA), 2021

  66. [74]

    Zhang, A

    Z. Zhang, A. Liniger, D. Dai, F. Yu, and L. V . Gool. End-to-end urban driving by imitating a reinforce- ment learning coach. arXiv.org, 2108.08265v3, 2021

  67. [75]

    P. Wu, X. Jia, L. Chen, J. Yan, H. Li, and Y . Qiao. Trajectory-guided control prediction for end-to- end autonomous driving: A simple yet strong baseline. In Advances in Neural Information Processing Systems (NeurIPS), 2022

  68. [76]

    X. Jia, P. Wu, L. Chen, J. Xie, C. He, J. Yan, and H. Li. Think twice before driving: Towards scalable decoders for end-to-end autonomous driving. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2023

  69. [77]

    X. Jia, Y . Gao, L. Chen, J. Yan, P. L. Liu, and H. Li. Driveadapter: Breaking the coupling barrier of perception and planning in end-to-end autonomous driving. In Proc. of the IEEE International Conf. on Computer Vision (ICCV), 2023

  70. [78]

    Ishida, G

    S. Ishida, G. Corrado, G. Fedoseev, H. Yeo, L. Russell, J. Shotton, J. F. Henriques, and A. Hu. Langprop: A code optimization framework using large language models applied to driving. InICLR 2024 Workshop on Large Language Model (LLM) Agents, 2024

  71. [79]

    Huang, R

    S. Huang, R. F. J. Dossa, C. Ye, J. Braga, D. Chakraborty, K. Mehta, and J. G. M. Araújo. Cleanrl: High-quality single-file implementations of deep reinforcement learning algorithms.Journal of Machine Learning Research (JMLR), 2022

  72. [80]

    Ha and J

    D. Ha and J. Schmidhuber. Recurrent world models facilitate policy evolution. In Advances in Neural Information Processing Systems (NeurIPS), 2018

  73. [81]

    Hafner, T

    D. Hafner, T. P. Lillicrap, J. Ba, and M. Norouzi. Dream to control: Learning behaviors by latent imagination. In Proc. of the International Conf. on Learning Representations (ICLR), 2020

  74. [82]

    Hafner, T

    D. Hafner, T. P. Lillicrap, M. Norouzi, and J. Ba. Mastering atari with discrete world models. In Proc. of the International Conf. on Learning Representations (ICLR), 2021

  75. [83]

    Kazemkhani, A

    S. Kazemkhani, A. Pandya, D. Cornelisse, B. Shacklett, and E. Vinitsky. Gpudrive: Data-driven, multi- agent driving simulation at 1 million FPS. In Proc. of the International Conf. on Learning Representa- tions (ICLR), 2025

  76. [84]

    Chen and P

    D. Chen and P. Krähenbühl. Learning from all vehicles. In Proc. IEEE Conf. on Computer Vision and Pattern Recognition (CVPR), 2022

  77. [85]

    J. Weng, M. Lin, S. Huang, B. Liu, D. Makoviichuk, V . Makoviychuk, Z. Liu, Y . Song, T. Luo, Y . Jiang, Z. Xu, and S. Yan. Envpool: A highly parallel reinforcement learning environment execution engine. In Advances in Neural Information Processing Systems (NeurIPS), 2022

  78. [86]

    Vinyals, I

    O. Vinyals, I. Babuschkin, W. M. Czarnecki, M. Mathieu, A. Dudzik, J. Chung, D. H. Choi, R. Powell, T. Ewalds, P. Georgiev, J. Oh, D. Horgan, M. Kroiss, I. Danihelka, A. Huang, L. Sifre, T. Cai, J. P. Agapiou, M. Jaderberg, A. S. Vezhnevets, R. Leblond, T. Pohlen, V . Dalibard...

  79. [87]

    Berner, G

    C. Berner, G. Brockman, B. Chan, V . Cheung, P. Debiak, C. Dennison, D. Farhi, Q. Fischer, S. Hashme, C. Hesse, R. Józefowicz, S. Gray, C. Olsson, J. Pachocki, M. Petrov, H. P. de Oliveira Pinto, J. Raiman, T. Salimans, J. Schlatter, J. Schneider, S. Sidor, I. Sutskever, J. Ta...

  80. [88]

    Petrenko, Z

    A. Petrenko, Z. Huang, T. Kumar, G. S. Sukhatme, and V . Koltun. Sample factory: Egocentric 3d control from pixels at 100000 FPS with asynchronous reinforcement learning. InProc. of the International Conf. on Machine learning (ICML), 2020

  81. [89]

    Pep 703 – making the global interpreter lock optional in cpython

    Python. Pep 703 – making the global interpreter lock optional in cpython. URL : https://peps.python.org/ pep-0703/, 2025

  82. [90]

    Paszke, S

    A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Z. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and S. Chintala. Pytorch: An imperative style, high-...

  83. [91]

    C. Lyle, Z. Zheng, E. Nikishin, B. Á. Pires, R. Pascanu, and W. Dabney. Understanding plasticity in neural networks. In Proc. of the International Conf. on Machine learning (ICML) , Proceedings of Machine Learning Research, 2023

  84. [92]

    Dohare, J

    S. Dohare, J. F. Hernandez-Garcia, Q. Lan, P. Rahman, A. R. Mahmood, and R. S. Sutton. Loss of plasticity in deep continual learning. Nature, 2024

  85. [93]

    Juliani and J

    A. Juliani and J. Ash. A study of plasticity loss in on-policy deep reinforcement learning. In Advances in Neural Information Processing Systems (NeurIPS), 2024

  86. [94]

    Moalla, A

    S. Moalla, A. Miele, D. Pyatko, R. Pascanu, and C. Gulcehre. No representation, no trust: Connecting representation, collapse, and trust issues in PPO. InAdvances in Neural Information Processing Systems (NeurIPS), 2024

  87. [95]

    P. Chou, D. Maturana, and S. A. Scherer. Improving stochastic policy gradients in continuous control with deep reinforcement learning using the beta distribution. In Proc. of the International Conf. on Machine learning (ICML), 2017

  88. [96]

    I. G. B. Petrazzini and E. A. Antonelo. Proximal policy optimization with continuous bounded action space via the beta distribution. In IEEE Symposium Series on Computational Intelligence, (SSCI), 2021

  89. [97]

    Jaeger, K

    B. Jaeger, K. Chitta, and A. Geiger. Hidden biases of end-to-end driving models. In Proc. of the IEEE International Conf. on Computer Vision (ICCV), 2023

  90. [98]

    Brockman, V

    G. Brockman, V . Cheung, L. Pettersson, J. Schneider, J. Schulman, J. Tang, and W. Zaremba. Openai gym. arXiv.org, 1606.01540, 2016

  91. [99]

    Towers, A

    M. Towers, A. Kwiatkowski, J. K. Terry, J. U. Balis, G. D. Cola, T. Deleu, M. Goulão, A. Kallinteris, M. Krimmel, A. KG, R. Perez-Vicente, A. Pierré, S. Schulhoff, J. J. Tai, H. Tan, and O. G. Younis. Gymnasium: A standard interface for reinforcement learning environments. arX...

  92. [100]

    Hintjens

    P. Hintjens. ZeroMQ: Messaging for Many Applications. O’Reilly Media, 2013

  93. [101]

    Huang, R

    S. Huang, R. F. J. Dossa, A. Raffin, A. Kanervisto, and W. Wang. The 37 implementation details of proximal policy optimization. In ICLR Blog Track, 2022. URL https://iclr-blog-track.github.io/2022/ 03/25/ppo-implementation-details/

  94. [102]

    Hanselmann, K

    N. Hanselmann, K. Renz, K. Chitta, A. Bhattacharyya, and A. Geiger. King: Generating safety-critical driving scenarios for robust imitation via kinematics gradients. In Proc. of the European Conf. on Com- puter Vision (ECCV), 2022

  95. [103]

    K. Hao, W. Cui, Y . Luo, L. Xie, Y . Bai, J. Yang, S. Yan, Y . Pan, and Z. Yang. Adversarial safety-critical scenario generation using naturalistic human driving priors. IEEE Trans. on Intelligent Transportation Systems (TITS), 2023

  96. [104]

    Y . Yin, P. Khayatan, Éloi Zablocki, A. Boulch, , and M. Cord. Regents: Real-world safety-critical driving scenario generation made stable. In ECCV 2024 W-CODA Workshop, 2024

  97. [105]

    D. R. Hipp. Sqlite. URL : http://sqlite.org/, 2000. 14

  98. [106]

    J. V . den Bossche, K. Jordahl, M. Fleischmann, M. Richards, J. McBride, J. Wasserman, A. G. Badaracco, A. D. Snow, B. Ward, J. Tratner, J. Gerard, M. Perry, cjqf, G. A. Hjelle, M. Taves, E. ter Hoeven, M. Cochran, R. Bell, rraymondgh, M. Bartos, P. Roggemans, L. Culbertson, G...

  99. [107]

    Zimmerlin, J

    J. Zimmerlin, J. Beißwenger, B. Jaeger, A. Geiger, and K. Chitta. Hidden biases of end-to-end driving datasets. arXiv.org, 2412.09602, 2024

  100. [108]

    Hafner, J

    D. Hafner, J. Pasukonis, J. Ba, and T. Lillicrap. Mastering diverse domains through world models. arXiv.org, 2023

  101. [109]

    Scheel, L

    O. Scheel, L. Bergamini, M. Wolczyk, B. Osi ´nski, and P. Ondruska. Urban driver: Learning to drive from real-world demonstrations using policy gradients. InProc. Conf. on Robot Learning (CoRL), 2021. 15 CaRL: Learning Scalable Planning Policies with Simple Rewards Supplementa...

  102. [110]

    official

    is a 3D open source autonomous driving simulator, based on the Unreal Engine, enabling real- time graphical rendering and simulation. CARLA enables research on the full driving stack from perception to planning and control, and has an extensive set of functionalities extended ...

Pith tools

Reviewed August 16, 2026 · model on record in the stance chip above.