REVIEW 19 cited by
Wayformer: Motion Forecasting via Simple & Efficient Attention Networks
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Wayformer: Motion Forecasting via Simple & Efficient Attention Networks
read the original abstract
Motion forecasting for autonomous driving is a challenging task because complex driving scenarios result in a heterogeneous mix of static and dynamic inputs. It is an open problem how best to represent and fuse information about road geometry, lane connectivity, time-varying traffic light state, and history of a dynamic set of agents and their interactions into an effective encoding. To model this diverse set of input features, many approaches proposed to design an equally complex system with a diverse set of modality specific modules. This results in systems that are difficult to scale, extend, or tune in rigorous ways to trade off quality and efficiency. In this paper, we present Wayformer, a family of attention based architectures for motion forecasting that are simple and homogeneous. Wayformer offers a compact model description consisting of an attention based scene encoder and a decoder. In the scene encoder we study the choice of early, late and hierarchical fusion of the input modalities. For each fusion type we explore strategies to tradeoff efficiency and quality via factorized attention or latent query attention. We show that early fusion, despite its simplicity of construction, is not only modality agnostic but also achieves state-of-the-art results on both Waymo Open MotionDataset (WOMD) and Argoverse leaderboards, demonstrating the effectiveness of our design philosophy
Forward citations
Cited by 19 Pith papers
-
SCORP: Scene-Consistent Multi-agent Diffusion Planning with Stable Online Reinforcement Post-Training for Cooperative Driving
SCORP delivers 10-28% gains in safety and 2-7% in efficiency metrics on WOMD by using dual-path scene conditioning in diffusion planning plus variance-gated group-relative policy optimization for closed-loop stability.
-
Class-Incremental Motion Forecasting
OMEN is the first end-to-end class-incremental motion forecaster that retains old-class accuracy via VLM-filtered future-detection pseudo-labels and variance-based sequence replay.
-
TRACER: Training-Free Closed-Loop Structured Inference for Traffic Accident Reconstruction
TRACER presents a training-free closed-loop structured inference framework for recovering physically consistent vehicle motions from sparse accident evidence.
-
SHIELD: Scalable Optimal Control with Certification using Duality and Convexity
SHIELD reduces decision variables and constraints in ℓ1-regularized convex programs via duality-derived certificates and transformer guidance, achieving order-of-magnitude speedups in stochastic MPC for traffic while ...
-
Zero-Shot Signal Temporal Logic Planning with Disjunctive Branch Selection in Dynamic Semantic Maps
A zero-shot STL planner combines a map-conditioned Transformer with a disjunctive heuristic and Transitive RL to achieve better generalization across dynamic semantic maps.
-
Learning physically grounded traffic accident reconstruction from public accident reports
A multimodal learning model with a new dataset of 6,217 cases reconstructs lane-consistent pre-impact motion and collision interactions from public accident reports, outperforming baselines in accuracy and consistency.
-
FlowS: One-Step Motion Prediction via Local Transport Conditioning
FlowS achieves state-of-the-art single-step motion prediction on Waymo Open Motion Dataset by using scene-conditioned anchor trajectories and a step-consistent displacement field to make local transport accurate in on...
-
EdgeVTP: Exploration of Latency-efficient Trajectory Prediction for Edge-based Embedded Vision Applications
EdgeVTP delivers the lowest measured end-to-end latency on Jetson-class platforms while matching or exceeding state-of-the-art accuracy on highway trajectory benchmarks by using bounded graph interactions and a one-sh...
-
SCORP: Scene-Consistent Multi-agent Diffusion Planning with Stable Online Reinforcement Post-Training for Cooperative Driving
Multi-ORFT improves closed-loop multi-agent driving planners by coupling scene-consistent diffusion pre-training with stable online RL post-training, reducing collisions and off-road rates while increasing speed on th...
-
A Dynamic Scene Interaction Reasoning Framework for Scene-level Lane-Change Intention and Trajectory Prediction of Multiple Interacting Vehicles
A dynamic graph-attention model jointly predicts every nearby vehicle’s lane-change intention and trajectory, cutting trajectory error by up to ~53% and improving scene coherence on NGSIM and highD.
-
Learning High-Level Decision Making with an Interaction-Aware Attention-Based Network in Autonomous Driving
An attention architecture that bottlenecks traffic agents into fixed latent queries plus a finer discrete action set yields higher simulated speeds and lower early-termination rates than DeepSet and Ego-attention on t...
-
Rethinking Training & Inference for Forecasting: Linking Winner-Take-All back to GMMs
Links WTA training mismatch in GMM-modeled forecasters to uninformative posteriors and introduces post-hoc merging plus one-step EM to yield better-ranked mode probabilities without retraining.
-
Learning A Unified Risk Map for Autonomous Driving in Partially Observable Environments
The paper proposes a unified risk map modeling and learning framework integrated with diffusion-based adversarial scenario generation for risk-aware planning in partially observable autonomous driving, demonstrating i...
-
Causality-Aware End-to-End Autonomous Driving via Ego-Centric Joint Scene Modeling
CaAD adds ego-centric joint-causal modeling and causality-aware policy alignment to end-to-end driving, reporting Driving Score 87.53 and Success Rate 71.81 on Bench2Drive plus PDMS 91.1 on NAVSIM.
-
Causality-Aware End-to-End Autonomous Driving via Ego-Centric Joint Scene Modeling
CaAD adds ego-centric joint-causal modeling and causality-aware policy alignment to end-to-end driving, reporting Driving Score 87.53 and PDMS 91.1 on Bench2Drive and NAVSIM.
-
SHIELD: Scalable Optimal Control with Certification using Duality and Convexity
SHIELD derives safe certificates from Lagrangian duality to reduce decision variables and constraints in convex programs, accelerated by a transformer network, delivering order-of-magnitude speedups in stochastic MPC ...
-
Recall to Predict: Grounding Motion Forecasting in Interpretable Motion Bank
A differentiable motion forecasting model retrieves and refines interpretable trajectory anchors from a contrastively learned motion bank to improve transparency without sacrificing multi-modal accuracy.
-
Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI
A survey of physical AI that distinguishes theoretical physics reasoning from applied understanding and synthesizes advances in symbolic reasoning, embodied systems, and generative models to advocate for physics-groun...
-
Transformer-Based Autonomous Driving Models and Deployment-Oriented Compression: A Survey
A survey that organizes Transformer-based autonomous driving models by task and architecture while analyzing compression techniques as a system-level deployment concern.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.