Pith. sign in

REVIEW 2 cited by

Real-time system optimal traffic routing under uncertainties -- Can physics models boost reinforcement learning?

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2407.07364 v1 pith:ZZL6ZA5Z submitted 2024-07-10 cs.LG cs.AIcs.SYeess.SY

classification cs.LGcs.AIcs.SYeess.SY
keywords learningtransrlphysicsmodelspolicyreinforcementsysteminterpretability
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

System optimal traffic routing can mitigate congestion by assigning routes for a portion of vehicles so that the total travel time of all vehicles in the transportation system can be reduced. However, achieving real-time optimal routing poses challenges due to uncertain demands and unknown system dynamics, particularly in expansive transportation networks. While physics model-based methods are sensitive to uncertainties and model mismatches, model-free reinforcement learning struggles with learning inefficiencies and interpretability issues. Our paper presents TransRL, a novel algorithm that integrates reinforcement learning with physics models for enhanced performance, reliability, and interpretability. TransRL begins by establishing a deterministic policy grounded in physics models, from which it learns from and is guided by a differentiable and stochastic teacher policy. During training, TransRL aims to maximize cumulative rewards while minimizing the Kullback Leibler (KL) divergence between the current policy and the teacher policy. This approach enables TransRL to simultaneously leverage interactions with the environment and insights from physics models. We conduct experiments on three transportation networks with up to hundreds of links. The results demonstrate TransRL's superiority over traffic model-based methods for being adaptive and learning from the actual network data. By leveraging the information from physics models, TransRL consistently outperforms state-of-the-art reinforcement learning algorithms such as proximal policy optimization (PPO) and soft actor critic (SAC). Moreover, TransRL's actions exhibit higher reliability and interpretability compared to baseline reinforcement learning approaches like PPO and SAC.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 2 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Know Unreported Roadway Incidents in Real-time: Early Traffic Anomaly Detection

    cs.LG 2024-12 reject novelty 5.0 of 10

    The paper proposes a deep learning framework that labels traffic anomalies from slowdown-speed patterns and trains sequence models to predict them up to 30 minutes ahead, reporting earlier alerts than incident-report-...

  2. SafeAug: Safety-Critical Driving Data Augmentation from Naturalistic Datasets

    cs.CV 2025-01 reject novelty 4.0 of 10

    A depth-based geometric augmentation pipeline moves the detected front vehicle closer in real KITTI images and rescales acceleration labels, reported to improve downstream emergency-braking prediction.

Pith tools