REVIEW 5 cited by
PlanT: Explainable Planning Transformers via Object-Level Representations
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
Planning an optimal route in a complex environment requires efficient reasoning about the surrounding scene. While human drivers prioritize important objects and ignore details not relevant to the decision, learning-based planners typically extract features from dense, high-dimensional grid representations containing all vehicle and road context information. In this paper, we propose PlanT, a novel approach for planning in the context of self-driving that uses a standard transformer architecture. PlanT is based on imitation learning with a compact object-level input representation. On the Longest6 benchmark for CARLA, PlanT outperforms all prior methods (matching the driving score of the expert) while being 5.3x faster than equivalent pixel-based planning baselines during inference. Combining PlanT with an off-the-shelf perception module provides a sensor-based driving system that is more than 10 points better in terms of driving score than the existing state of the art. Furthermore, we propose an evaluation protocol to quantify the ability of planners to identify relevant objects, providing insights regarding their decision-making. Our results indicate that PlanT can focus on the most relevant object in the scene, even when this object is geometrically distant.
Forward citations
Cited by 5 Pith papers
-
Deconfounded Lifelong Learning for Autonomous Driving via Dynamic Knowledge Spaces
DeLL combines DPMM dual knowledge spaces with front-door causal adjustment and a non-autoregressive evolutionary decoder to reduce catastrophic forgetting and spurious correlations in lifelong end-to-end autonomous driving.
-
DiffE2E: Rethinking End-to-End Driving with a Hybrid Action Diffusion and Supervised Policy
DiffE2E reports state-of-the-art closed-loop driving scores in CARLA and NAVSIM by combining a diffusion trajectory decoder with explicit supervision in a single Transformer decoder.
-
Agent-driven Long-tail Simulation for Autonomous Driving
LLM agents with structured actions can drive interactive long-tail road users in nuPlan, and SemanticPlan shows current planners still fail safety and semantic completion there.
-
Large Language Model Enhanced Differentiable Trajectory Planning for IoT-Enabled Autonomous Driving
An IL planner with agent-centric data reuse, complexity-aware async LLM semantics, and residual differentiable optimization reports top nuPlan Hard20 closed-loop scores and real-time CARLA-ROS execution.
-
A Survey on Vision-Language-Action Models for Autonomous Driving
A survey organizes vision-language-action models for autonomous driving into four stages, compares over 20 systems, and catalogs datasets, benchmarks, and open challenges.
Discussion (0). Sign in to comment.