REVIEW 4 major objections 4 minor 45 references
This paper claims that a minimum-lap-time trajectory computed by optimal control can replace expert demonstrations and guide a staged reinforcement-learning controller to execute autonomous drifting faster than a human driver in simulation.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-04 01:15 UTC pith:QX5EOW76
load-bearing objection A promising integration of MLT planning and curriculum RL for drifting, but internal inconsistencies and an untested planner-to-simulator gap leave the central claim unsupported as written. the 4 major comments →
Track-Guided Hierarchical Reinforcement Learning for Autonomous Vehicle Drifting with Minimum-Lap-Time Planning
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
The central claim is that minimum-lap-time drift planning (MLTDP) provides a good enough reference that a track-guided RL controller (TgRL) can learn to stabilize large-sideslip drifting and reduce lap time. The paper formulates drifting as an optimal control problem in Frenet coordinates, minimizing travel time over arc length subject to vehicle dynamics, tire limits, road boundaries, and state and control constraints. The resulting reference trajectory supplies speed, sideslip angle, and yaw-rate targets. An entropy-regularized actor-critic agent is then trained through a three-stage curriculum—basic drift control, drift cornering, drift racing—with a reward that combines tracking errors a
What carries the argument
The load-bearing objects are the minimum-lap-time reference trajectory and the three-stage curriculum. The planner solves an optimal control problem over arc length, minimizing the integral of (1-l·κ)/(v·cos(Δψ+β)) under a 3-DoF bicycle model with Magic Formula tire forces, producing positions, speeds, sideslip angles, and yaw rates that serve as prior data. The controller then uses that trajectory in stagewise training: first learning to hold a drift on a straight track, then applying it through hairpins, then combining both into a full-lap race policy. The reward's terminal term is scaled by the gap between achieved lap time and the planned optimal lap time, which is what connects the RL o
Load-bearing premise
The plan's value hinges on the assumption that the simplified vehicle-and-tire model, tuned for a low-friction road, produces a trajectory the simulated car can actually follow and that is close to the fastest possible trajectory for that car.
What would settle it
Retrain the controller with the planner's reference replaced by a randomly perturbed or centerline trajectory while keeping the reward structure fixed; if lap times stay similar, the minimum-lap-time prior is not the cause of the speed-up. Alternatively, measure tracking of the planned sideslip profile at the two corners where the paper reports the largest deviations; if the agent cannot follow the drift posture there, the lap-time gain would have to come from something other than executing the MLT plan.
If this is right
- An RL drift controller can be trained without human expert demonstrations, using an optimal-control trajectory as the prior.
- A three-stage curriculum (straight-line drift, corner drift, full-lap race) makes learning feasible where a plain actor-critic agent gets trapped in negative reward.
- The trained policy generalizes across three different track layouts and consistently beats a prior RL drifting baseline, with the kart-track average lap time dropping from 51.2 s to 45.4 s.
- Executed drifts reach average sideslip angles around 16–18 degrees, well beyond the steady-state cornering envelope, while keeping yaw error lower than the baseline.
- The framework approaches the idealized planner's lap time within about 12–17 percent on the tested tracks.
Where Pith is reading between the lines
- Our inference: if the planner-to-simulator model gap is the main source of the remaining 12–17 percent, updating the plan online with a learned dynamics model could close much of that gap without changing the curriculum.
- Our inference: the action space restricts throttle to 0.6–1.0 and omits braking, so the controller cannot reproduce planner behavior that requires lifting off the throttle or braking; part of the measured gap may stem from this restriction rather than from policy failure.
- Our inference: the same recipe—optimal-control trajectory as a curriculum prior plus a lap-time-scaled terminal reward—should transfer to other extreme maneuvers, such as drift parking or emergency obstacle avoidance, where a planner can supply the reference and RL absorbs model error.
- Our inference: a sharper test of the mechanism would corrupt only the planner's sideslip profile while keeping the geometric path fixed; if lap times barely change, the benefit comes from the racing line itself rather than from learning the drift posture.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. This paper proposes MLTDP-TgRL, a hierarchical planning/RL framework for autonomous drifting. MLTDP formulates a minimum-lap-time optimal control problem in Frenet coordinates using a 3-DoF bicycle model with a Magic Formula tire model (μ=0.6). The resulting trajectory is used as a reference prior for a SAC-based controller trained in CARLA via a three-stage curriculum: basic drift control, corner drifting, and full race policy. The reward combines instant tracking terms with an end reward based on the planner's lap time. Evaluations on Lkarting, Willow Springs, and Nelson Ledges compare against a published drift RL baseline (Cai_DRIFT) and a human driver; ablations vary the reference (MLT vs centerline), RL variant, end reward, and action smoothing. The paper claims superior lap times and larger sideslip angles while maintaining tracking accuracy.
Significance. If the claims hold, the framework is a useful step toward replacing expert demonstrations with model-based trajectory priors for drift control, and the three-stage curriculum plus action smoothing are sensible engineering contributions. The comparison to a published drift RL baseline and human data on three tracks, together with the ablation studies, is a substantial evaluation. However, the current manuscript does not yet substantiate the central lap-time claim because of numerical inconsistencies in key results, missing uncertainty quantification, and incomplete validation of planner feasibility. These issues are fixable but require a revision.
major comments (4)
- [§5.4, Table 4] The headline Lkarting result is internally inconsistent: the text says MLT+TgRL "completes laps in an average of 55.4 s," but Table 4 gives Avg = 45.4 s for MLT+TgRL and 51.2 s for Cai_DRIFT. The stated "approximately 10%" gain is consistent with 45.4 vs 51.2, so 55.4 appears to be a typo, but it is central and must be corrected. Similarly, Sec. 5.3 reports TgRL-Control lap time "about 36.5 s," while Table 4 lists Min = 37.2 s. These numbers need to be reconciled before the results can be assessed.
- [§4.4.2, Table 3, §5.3] The end reward is claimed to be derived from the MLTDP optimum, but the illustrative example uses Ttotal = 46 s and k6 = 5e5, while Sec. 5.3 and Table 4 report MLT Planning ≈33 s/33.1 s and Table 3 lists k6 = 1e5. With Ttotal = 33.1 s and k6 = 1e5, a 50 s lap would give an end reward of about 3.4e3, not 2.25e5. Because the end reward is one of the key mechanisms for transferring the MLT objective into RL, the actual Ttotal used for Lkarting (and per track, if varied) must be reported consistently; otherwise the claim that the reward encodes the minimum-lap-time objective is not verifiable.
- [§5.3, §5.4] The MLTDP trajectory is computed from a simplified 3-DoF bicycle/Magic Formula model (μ=0.6) but evaluated with CARLA's built-in vehicle model. The paper acknowledges "differences in the vehicle models used for planning and control" and reports a large sideslip gap (~30° planned vs ~18° executed). Yet contribution (2) calls these "dynamically feasible" trajectories and Sec. 5.4 treats MLT Planning as a "theoretical lower bound." No experiment shows that the planner's solution is dynamically feasible for the CARLA vehicle. Since the main claim is that the MLT prior, not just the reward shaping/curriculum, produces the lap-time gains, the paper should either (i) report closed-loop tracking of the MLT trajectory (e.g., with a well-tuned MPC or an oracle tracker), or (ii) soften the lower-bound language and state explicitly that the planner is a heuristic prior. Figure 12(a)'s MLT-vs-centerl
- [Table 4, §5.4, Fig. 12] There is no uncertainty quantification. The averages in Table 4 are over ten laps, but no standard deviations, confidence intervals, number of training seeds, or statistical comparisons are reported. RL training is stochastic, and the differences against Cai_DRIFT on Willow Springs (Min 181.1 vs 193.1; Avg 197.5 vs 197.7) are small enough that seed-to-seed variation could change the conclusion. Similarly, Fig. 12 reports single numbers per method. Add per-seed results, error bars, and an explicit statement of how many seeds and laps were used; otherwise the claim of "consistently" superior performance is not supported.
minor comments (4)
- [Fig. 12(e), §4.3] The text calls AAS an "Adaptive Angle Shaping reward," but in Sec. 4.3 AAS is Adaptive Action Smoothing, not a reward. Correct the label and terminology.
- [Table 3, Eq. (23)] The ASS speed range is listed as [16.67, 20] km/h, but vehicle speeds in the paper are in m/s (e.g., 0–30 m/s in Eq. 16). If the intended values are m/s, the range is implausibly low for the stated purpose; if km/h, convert to m/s and use consistent units.
- [§5.4] The sentence that the Lkarting average sideslip angle of 15.9° is "close to the optimal planning reference of 29.8°" is misleading; 15.9° is roughly half of 29.8°. Clarify that 29.8° is a per-lap maximum, not a time-average, or rephrase.
- [Table 4] The column header "Sideslip Angle Avg (Laps)" is confusing, and the MLT Planning rows present a single run with no indication of repeatability. Define the metric precisely and use consistent formatting.
Circularity Check
No significant circularity: lap-time results are measured against external baselines; MLTDP trajectory is an independent OCP output used only as reward guidance.
full rationale
The paper's derivation chain is not circular in the sense that matters here. The MLTDP trajectory is produced by solving the optimal-control problem in Eqs. (7)-(15) from the 3-DoF bicycle/Magic Formula model, with no dependence on the RL controller or on the reported lap times. The RL agent is then trained with a reward (Eqs. 24-27) that tracks this trajectory and uses T_total from MLTDP as an end-reward anchor. This is intentional guidance, not a fitted parameter renamed as a prediction: the final lap times in Table 4 are measured in CARLA against external baselines (Cai_DRIFT and a human driver), so the claimed lap-time improvement is empirical and falsifiable. The ablation study (Fig. 12) also compares against centerline-tracking agents, providing independent evidence that the MLT reference, not merely the reward shaping, contributes. The Sec. 5.3 admission of vehicle-model mismatch ('This discrepancy is likely due to differences in the vehicle models used for planning and control') is a feasibility/correctness concern, not a circular step: the planner is not defined in terms of the RL outcome. The inconsistency between the Sec. 4.4.2 example T_total=46 s and the Sec. 5.3/Table 4 MLT Planning time of ~33 s weakens the claim that the end reward is faithfully tied to the planner's optimum, but it does not make the derivation reduce to its inputs. No load-bearing self-citations, uniqueness theorems, or ansatz-smuggling are present; the self-citations are background. Therefore the circularity score is 0.
Axiom & Free-Parameter Ledger
free parameters (8)
- Magic Formula tire coefficients B,C,D,E =
not provided
- Reward function kernel widths k1,k2,k3 =
8, 3200, 5
- Instant reward weights k_ey,k_eψ,k_eβ,k_ev =
20, 40, 20, 40
- Velocity discount K and threshold v_thre =
0.5, 5 m/s
- End reward coefficients k4,k5,k6,k7 =
1e5, -0.05, 1e5, -0.2 (Table 3; Sec 4.4.2 example uses k6=5e5)
- AAS smoothing parameters λτ, λδ,min, λδ,max, [v_min,v_max] =
λτ=0.2; [16.67,20] km/h; λδ range not given
- Lookahead horizon N =
10
- Curriculum stage episode counts M,P =
not specified
axioms (6)
- domain assumption The 3-DoF bicycle model (Eq 1) and Magic Formula combined-slip tire model (Eq 5-6) adequately describe vehicle drift dynamics for planning.
- domain assumption The road friction coefficient is constant at μ=0.6 on all three tracks (Table 1).
- domain assumption The CARLA simulator's built-in vehicle model is a valid testbed that captures the nonlinearities of drifting.
- domain assumption SNOPT converges to a sufficiently accurate local solution of the MLT OCP.
- ad hoc to paper The hand-designed reward (Eq 24-28) correctly encodes the minimum-lap-time objective.
- ad hoc to paper The three-stage curriculum transfers to full-track racing without simulator-to-simulator gaps beyond the reported transfer shocks.
read the original abstract
In Formula 1, drivers optimize racing lines within tire grip limits to minimize lap times; however, in rally racing, drivers intentionally break traction to drift on loose surfaces. This maneuver rapidly aligns the vehicle for corner exits, ultimately reducing lap time. Autonomously executing such maneuvers formulates a complex dual-objective control problem: stabilizing highly nonlinear drift dynamics while strictly minimizing lap time. Addressing this challenge motivates the development of advanced Minimum-Lap-Time (MLT) drift control architectures. This paper proposes a planning-control framework specifically designed for MLT drifting scenario. First, we formulate an optimal control problem to generate a MLT drift planning trajectory, which is used as prior data to train a deep reinforcement learning drift controller. Given that drifting involves extremely large sideslip angles and is therefore challenging to learn directly, a Track-guided Reinforcement Learning (TgRL) drift control method is proposed to enable progressive training in a step-by-step manner, from drift control policy, to drift corner policy, and finally to a comprehensive drift race policy. The reward function incorporates both an instant reward term and an end reward term derived from the Minimum-Lap-Time objective. Simulation results demonstrate that the proposed framework enables the agent to learn a drift racing policy that not only ensures vehicle motion control performance but also effectively reduces lap time.
Figures
Reference graph
Works this paper leans on
-
[1]
R. Y. Hindiyeh, J. Christian Gerdes, A controller framework for autonomous drifting: Design, stability, and experimental validation, J. of Dyn. Syst., Meas., and Control 136 (2014) 051015
2014
-
[2]
Z. Shan, J. Zhao, B. Zhu, C. Lv, Y. Zhao, L. Ge, S. Zhong, Safe and efficient trajectory planning considering longitudinal and lateral coupled limits, IEEE Trans. Veh. Technol. 73 (2024) 10714–10719
2024
-
[3]
J. Ni, J. Hu, C. Xiang, Envelope control for four-wheel independently actuated autonomous ground vehicle through afs/dyc integrated con- trol, IEEE Trans. Veh. Technol. 66 (2017) 9712–9726
2017
-
[4]
Cheng, B.-B
S. Cheng, B.-B. Hu, H.-L. Wei, L. Li, C. Lv, Deep learning-based hybrid dynamic modeling and improved handling stability assessment for autonomous vehicles at driving limits, IEEE Trans. Veh. Technol. 74 (2025) 5582–5593
2025
-
[5]
N. A. Spielberg, M. Brown, N. R. Kapania, J. C. Kegelman, J. C. Gerdes, Neural network vehicle models for high-performance auto- mated driving, Sci. Robot. 4 (2019) eaaw1975
2019
-
[6]
N. D. Broadbent, T. Weber, D. Mori, J. C. Gerdes, Neural network tire force modeling forăautomated drifting, in: 16th Int. Symp. on Adv. Vehicle Control, Springer Nature Switzerland, Cham, 2024, pp. 378–384
2024
-
[7]
T. P. Weber, J. C. Gerdes, Modeling and control for dynamic drifting trajectories, IEEE Trans. Intell. Veh. 9 (2024) 3731–3741
2024
-
[8]
T. P. Weber, R. K. Aggarwal, J. C. Gerdes, Human-inspired au- tonomous racing in low friction environments, IEEE Trans. Intell. Veh. (2024) 1–14
2024
-
[9]
Q. Ma, X. Yin, X. Zhang, X. Xu, X. Yao, Game-theoretic receding- horizon reinforcement learning for lateral control of autonomous vehicles, IEEE Trans. Veh. Technol. 73 (2024) 14547–14562
2024
-
[10]
X. Wu, J. Li, C. Su, J. Fan, M. Xu, A deep reinforcement learning based hierarchical eco-driving strategy for connected and automated hevs, IEEE Trans. Veh. Technol. 72 (2023) 13901–13916
2023
-
[11]
S. Zhao, J. Zhang, N. Masoud, Y. Jiang, H. Huang, T. Liu, Drift cornering control and real-vehicle deployment for electric vehicles, IEEE Transactions on Industrial Electronics 72 (2025) 13509–13520
2025
-
[12]
P. R. Wurman, et al., Outracing champion gran turismo drivers with deep reinforcement learning, Nature 602 (2022) 223–228
2022
-
[13]
P. Cai, X. Mei, L. Tai, Y. Sun, M. Liu, High-speed autonomous drifting with deep reinforcement learning, IEEE Robot. and Automat. Lett. 5 (2020) 1247–1254
2020
- [14]
-
[15]
Y. Yin, S. E. Li, K. Li, J. Yang, F. Ma, Self-learning drift control of automated vehicles beyond handling limit after rear-end collision, Transp. Saf. and Environ. 2 (2020) 97–105
2020
-
[16]
S. Zhao, J. Zhang, C. He, X. Hou, H. Huang, Adaptive drift control of autonomous electric vehicles after brake system failures, IEEE Trans. Ind. Electron. 71 (2024) 6041–6052
2024
-
[17]
S. H. Tóth, Ádám Bárdos, Z. J. Viharos, Tabular Q-learning based reinforcement learning agent for autonomous vehicle drift initiation and stabilization, IFAC-PapersOnLine 56 (2023) 4896–4903. 22nd IFAC World Congress
2023
-
[18]
S. H. Tóth, Z. J. Viharos, . Bárdos, Z. Szalay, Sim-to-real application of reinforcement learning agents for autonomous, real vehicle drift- ing, Vehicles 6 (2024) 781–798
2024
-
[19]
Y. Wang, X. Yuan, C. Sun, Learning autonomous race driving with action mapping reinforcement learning, ISA Trans. 150 (2024) 1–14. Sheng Zhao et al.: Preprint submitted to Elsevier Page 14 of 15 Track-Guided Hierarchical Reinforcement Learning for Autonomous Vehicle Drifting with Minimum-Lap-Time Planning
2024
-
[20]
X. Hou, J. Zhang, C. He, Y. Ji, J. Zhang, J. Han, Autonomous driving at the handling limit using residual reinforcement learning, Adv. Eng. Inform. 54 (2022) 101754
2022
-
[21]
F. Domberg, C. C. Wembers, H. Patel, G. Schildbach, Deep drifting: Autonomous drifting of arbitrary trajectories using deep reinforce- ment learning, in: 2022 IEEE Int. Conf. on Robot. and Automat. (ICRA), 2022, pp. 7753–7759. doi: 10.1109/ICRA46639.2022.9812249
arXiv 2022
-
[22]
Domberg, B
F. Domberg, B. Barkow, G. Schildbach, Vision-based autonomous trajectory drifting using deep reinforcement learning, in: AmEC 2024 Automot. meets Electro. and Control; 14. GMM Symp., 2024, pp. 47– 52
2024
-
[23]
F. Djeumou, M. Thompson, M. Suminaka, J. Subosits, Reference-free formula drift with reinforcement learning: From driving data to tire energy-inspired, real-world policies, ArXiv abs/2410.20990 (2024)
Pith/arXiv arXiv 2024
-
[24]
Orgován, T
L. Orgován, T. Bécsi, S. Aradi, Autonomous drifting using rein- forcement learning, Periodica Polytechnica Transp. Eng. 49 (2021) 292300
2021
-
[25]
B. Leng, Y. Yu, M. Liu, et al., Deep reinforcement learning-based drift parking control of automated vehicles, Sci. China Technological Sciences 66 (2023) 1152–1165
2023
-
[26]
Bhattacharjee, D
S. Bhattacharjee, D. Schnieders, Autonomous drifting RC car with reinforcement learning, interim report (2018)
2018
-
[27]
Hoshino, J
H. Hoshino, J. Li, A. Menon, J. M. Dolan, Y. Nakahira, Autonomous drifting based on maximal safety probability learning, in: 2024 IEEE 27th Int. Conf. on Intell. Transp. Syst. (ITSC 2024), 2024, pp. 3930–
2024
-
[28]
J. Li, X. Wu, M. Xu, Y. Liu, Deep reinforcement learning and reward shaping based eco-driving control for automated hevs among signalized intersections, Energy 251 (2022). Cited by: 82
2022
-
[29]
J. Li, A. Fotouhi, W. Pan, Y. Liu, Y. Zhang, Z. Chen, Deep rein- forcement learning-based eco-driving control for connected electric vehicles at signalized intersections considering traffic uncertainties, Energy 279 (2023). Cited by: 42; All Open Access, Green Open Access
2023
-
[30]
D. Li, J. Zhang, S. Lin, Planning and control of drifting-based collision avoidance strategy under emergency driving conditions, Control Eng. Pract. 139 (2023) 105625
2023
-
[31]
A. Bertipaglia, D. Tavernini, U. Montanaro, M. Alirezaei, R. Happee, A. Sorniotti, B. Shyrokau, Model predictive contouring control for vehicle obstacle avoidance at the limit of handling using torque vectoring*, in: 2024 IEEE International Conference on Advanced Intelligent Mechatronics (AIM), 2024, pp. 1468–1475. doi: 10.1109/ AIM55361.2024.10637113
arXiv 2024
-
[32]
X. Zhao, G. Chen, Z. Gao, J. Yao, Z. Gao, M. Hua, Autonomous obstacle avoidance for distributed drive electric vehicles via dynamic drifting, IEEE Trans. on Transp. Electrific. 10 (2024) 8893–8906
2024
-
[33]
Kabzan, L
J. Kabzan, L. Hewing, A. Liniger, M. N. Zeilinger, Learning-based model predictive control for autonomous racing, IEEE Robot. and Automat. Lett. 4 (2019) 3363–3370
2019
-
[34]
Djeumou, T
F. Djeumou, T. J. Lew, N. Ding, M. Thompson, M. Suminaka, M. Greiff, J. Subosits, One model to drift them all: Physics-informed conditional diffusion model for driving at the limits, in: 8th Annual Conf. on Robot Learn., 2024
2024
-
[35]
Huang, Y
Y. Huang, Y. Chen, Vehicle lateral stability control based on shiftable stability regions and dynamic margins, IEEE Trans. Veh. Technol. 69 (2020) 14727–14738
2020
-
[36]
H. Lu, X. Wu, S. Zhao, L. Yan, J. Lu, Controlling nonlinear vehicular motions by exploiting linearized feedback law under delay-tolerance: stability, gain-scheduling, and validation, Meccanica (2025)
2025
-
[37]
Ajanovi, E
Z. Ajanovi, E. Regolin, B. Shyrokau, H. ati, M. Horn, A. Ferrara, Search-based task and motion planning for hybrid systems: Agile autonomous vehicles, Eng. Appl. of Artif. Intell. 121 (2023) 105893
2023
-
[38]
J. Li, X. Wu, X. Bai, Y. Liu, M. Xu, Intelligent eco-driving control for urban cavs using a model-based controller assisted deep reinforce- ment learning, IEEE Trans. Intell. Transp. Syst. 26 (2025) 7624–7639
2025
-
[39]
J. Li, X. Wu, J. Fan, Y. Liu, M. Xu, Overcoming driving challenges in complex urban traffic: A multi-objective eco-driving strategy via safety model based reinforcement learning, Energy 284 (2023). Cited by: 23
2023
-
[40]
J. Li, X. Wu, M. Xu, Y. Liu, Multiobjective eco-driving strategy for connected and automated electric vehicles considering complex urban traffic influence factors, IEEE Transactions on Transportation Electrification 10 (2024) 10043 10058. Cited by: 5
2024
-
[41]
Haarnoja, A
T. Haarnoja, A. Zhou, P. Abbeel, S. Levine, Soft actor-critic: Off- policy maximum entropy deep reinforcement learning with a stochas- tic actor, in: J. Dy, A. Krause (Eds.), Proc. of the 35th Int. Conf. on Mach. Learn., volume 80 of Proc. of Mach. Learn. Res., PMLR, 2018, pp. 1861–1870
2018
-
[42]
Huang, H
W. Huang, H. Liu, Z. Huang, C. Lv, Safety-aware human-in-the-loop reinforcement learning with shared control for autonomous driving, IEEE Trans. Intell. Transp. Syst. 25 (2024) 16181–16192
2024
-
[43]
P. E. Gill, W. Murray, M. A. Saunders, Snopt: An sqp algorithm for large-scale constrained optimization, SIAM Rev. 47 (2005) 99–131
2005
-
[44]
Dosovitskiy, G
A. Dosovitskiy, G. Ros, F. Codevilla, A. Lopez, V . Koltun, CARLA: An open urban driving simulator, in: S. Levine, V . Vanhoucke, K. Goldberg (Eds.), Proc.s of the 1st Annu. Conf. on Robot Learn., volume 78 of Proceedings of Mach. Learn. Res. , PMLR, 2017, pp. 1–16. Sheng Zhao et al.: Preprint submitted to Elsevier Page 15 of 15
2017
-
[3935]
doi:10.1109/ITSC58415.2024.10919509
arXiv 2024
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.