Pith. sign in

REVIEW 2 cited by

End-to-End Urban Driving by Imitating a Reinforcement Learning Coach

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2108.08265 v3 pith:FGAYBBHG submitted 2021-08-18 cs.CV cs.RO

classification cs.CVcs.RO
keywords end-to-enddrivinglearningcoachexpertperformancereinforcementachieves
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

End-to-end approaches to autonomous driving commonly rely on expert demonstrations. Although humans are good drivers, they are not good coaches for end-to-end algorithms that demand dense on-policy supervision. On the contrary, automated experts that leverage privileged information can efficiently generate large scale on-policy and off-policy demonstrations. However, existing automated experts for urban driving make heavy use of hand-crafted rules and perform suboptimally even on driving simulators, where ground-truth information is available. To address these issues, we train a reinforcement learning expert that maps bird's-eye view images to continuous low-level actions. While setting a new performance upper-bound on CARLA, our expert is also a better coach that provides informative supervision signals for imitation learning agents to learn from. Supervised by our reinforcement learning coach, a baseline end-to-end agent with monocular camera-input achieves expert-level performance. Our end-to-end agent achieves a 78% success rate while generalizing to a new town and new weather on the NoCrash-dense benchmark and state-of-the-art performance on the challenging public routes of the CARLA LeaderBoard.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. CaRL: Learning Scalable Planning Policies with Simple Rewards

    cs.LG 2025-04 accept novelty 7.0 of 10

    A route-completion reward with episode termination and multiplicative soft penalties enables PPO to scale to 300M CARLA and 500M nuPlan samples, reaching 64 DS on longest6 v2 and 91 CLS on Val14.

  2. A Comprehensive Review of Reinforcement Learning for Autonomous Driving in the CARLA Simulator

    cs.RO 2025-09 conditional novelty 4.0 of 10

    A survey of roughly 100 CARLA reinforcement learning papers, mapping algorithm families, representations, rewards, evaluation metrics, towns, and open challenges.

Pith tools