Pith. sign in

REVIEW 1 cited by

Extensive Exploration in Complex Traffic Scenarios using Hierarchical Reinforcement Learning

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2501.14992 v1 pith:O7AQ74GT submitted 2025-01-25 cs.LG cs.RO

classification cs.LGcs.RO
keywords controllercomplexdrivinghierarchicalrewardsscenariostrafficcontrollers
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Developing an automated driving system capable of navigating complex traffic environments remains a formidable challenge. Unlike rule-based or supervised learning-based methods, Deep Reinforcement Learning (DRL) based controllers eliminate the need for domain-specific knowledge and datasets, thus providing adaptability to various scenarios. Nonetheless, a common limitation of existing studies on DRL-based controllers is their focus on driving scenarios with simple traffic patterns, which hinders their capability to effectively handle complex driving environments with delayed, long-term rewards, thus compromising the generalizability of their findings. In response to these limitations, our research introduces a pioneering hierarchical framework that efficiently decomposes intricate decision-making problems into manageable and interpretable subtasks. We adopt a two step training process that trains the high-level controller and low-level controller separately. The high-level controller exhibits an enhanced exploration potential with long-term delayed rewards, and the low-level controller provides longitudinal and lateral control ability using short-term instantaneous rewards. Through simulation experiments, we demonstrate the superiority of our hierarchical controller in managing complex highway driving situations.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Bootstrapping Reinforcement Learning with Sub-optimal Policies for Autonomous Driving

    cs.RO 2025-09 conditional novelty 4.0 of 10

    A sub-optimal rule-based controller used as a soft constraint and replay-buffer data source lets a SAC agent escape a highway slow-traffic trap, outperforming SAC, CQL, and GAIL.

Pith tools