Pith. sign in

REVIEW 19 cited by

Wayformer: Motion Forecasting via Simple & Efficient Attention Networks

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2207.05844 v1 pith:GNC4AE76 submitted 2022-07-12 cs.CV

Wayformer: Motion Forecasting via Simple & Efficient Attention Networks

classification cs.CV
keywords attentionforecastingfusionmotionwayformercomplexdesigndiverse
verification ladder T0 review T1 audit T2 compute T3 formal T4 reserved
0 comments
read the original abstract

Motion forecasting for autonomous driving is a challenging task because complex driving scenarios result in a heterogeneous mix of static and dynamic inputs. It is an open problem how best to represent and fuse information about road geometry, lane connectivity, time-varying traffic light state, and history of a dynamic set of agents and their interactions into an effective encoding. To model this diverse set of input features, many approaches proposed to design an equally complex system with a diverse set of modality specific modules. This results in systems that are difficult to scale, extend, or tune in rigorous ways to trade off quality and efficiency. In this paper, we present Wayformer, a family of attention based architectures for motion forecasting that are simple and homogeneous. Wayformer offers a compact model description consisting of an attention based scene encoder and a decoder. In the scene encoder we study the choice of early, late and hierarchical fusion of the input modalities. For each fusion type we explore strategies to tradeoff efficiency and quality via factorized attention or latent query attention. We show that early fusion, despite its simplicity of construction, is not only modality agnostic but also achieves state-of-the-art results on both Waymo Open MotionDataset (WOMD) and Argoverse leaderboards, demonstrating the effectiveness of our design philosophy

discussion (0)

Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.

Forward citations

Cited by 19 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. SCORP: Scene-Consistent Multi-agent Diffusion Planning with Stable Online Reinforcement Post-Training for Cooperative Driving

    cs.RO 2026-04 unverdicted novelty 7.0

    SCORP delivers 10-28% gains in safety and 2-7% in efficiency metrics on WOMD by using dual-path scene conditioning in diffusion planning plus variance-gated group-relative policy optimization for closed-loop stability.

  2. Class-Incremental Motion Forecasting

    cs.CV 2026-03 conditional novelty 7.0

    OMEN is the first end-to-end class-incremental motion forecaster that retains old-class accuracy via VLM-filtered future-detection pseudo-labels and variance-based sequence replay.

  3. TRACER: Training-Free Closed-Loop Structured Inference for Traffic Accident Reconstruction

    cs.LG 2026-06 unverdicted novelty 6.0

    TRACER presents a training-free closed-loop structured inference framework for recovering physically consistent vehicle motions from sparse accident evidence.

  4. SHIELD: Scalable Optimal Control with Certification using Duality and Convexity

    cs.RO 2026-05 unverdicted novelty 6.0

    SHIELD reduces decision variables and constraints in ℓ1-regularized convex programs via duality-derived certificates and transformer guidance, achieving order-of-magnitude speedups in stochastic MPC for traffic while ...

  5. Zero-Shot Signal Temporal Logic Planning with Disjunctive Branch Selection in Dynamic Semantic Maps

    cs.AI 2026-05 unverdicted novelty 6.0

    A zero-shot STL planner combines a map-conditioned Transformer with a disjunctive heuristic and Transitive RL to achieve better generalization across dynamic semantic maps.

  6. Learning physically grounded traffic accident reconstruction from public accident reports

    cs.LG 2026-04 unverdicted novelty 6.0

    A multimodal learning model with a new dataset of 6,217 cases reconstructs lane-consistent pre-impact motion and collision interactions from public accident reports, outperforming baselines in accuracy and consistency.

  7. FlowS: One-Step Motion Prediction via Local Transport Conditioning

    cs.RO 2026-04 unverdicted novelty 6.0

    FlowS achieves state-of-the-art single-step motion prediction on Waymo Open Motion Dataset by using scene-conditioned anchor trajectories and a step-consistent displacement field to make local transport accurate in on...

  8. EdgeVTP: Exploration of Latency-efficient Trajectory Prediction for Edge-based Embedded Vision Applications

    cs.CV 2026-04 unverdicted novelty 6.0

    EdgeVTP delivers the lowest measured end-to-end latency on Jetson-class platforms while matching or exceeding state-of-the-art accuracy on highway trajectory benchmarks by using bounded graph interactions and a one-sh...

  9. SCORP: Scene-Consistent Multi-agent Diffusion Planning with Stable Online Reinforcement Post-Training for Cooperative Driving

    cs.RO 2026-04 unverdicted novelty 6.0

    Multi-ORFT improves closed-loop multi-agent driving planners by coupling scene-consistent diffusion pre-training with stable online RL post-training, reducing collisions and off-road rates while increasing speed on th...

  10. A Dynamic Scene Interaction Reasoning Framework for Scene-level Lane-Change Intention and Trajectory Prediction of Multiple Interacting Vehicles

    cs.AI 2026-07 conditional novelty 5.5

    A dynamic graph-attention model jointly predicts every nearby vehicle’s lane-change intention and trajectory, cutting trajectory error by up to ~53% and improving scene coherence on NGSIM and highD.

  11. Learning High-Level Decision Making with an Interaction-Aware Attention-Based Network in Autonomous Driving

    cs.RO 2026-06 conditional novelty 5.0

    An attention architecture that bottlenecks traffic agents into fixed latent queries plus a finer discrete action set yields higher simulated speeds and lower early-termination rates than DeepSet and Ego-attention on t...

  12. Rethinking Training & Inference for Forecasting: Linking Winner-Take-All back to GMMs

    cs.LG 2026-06 unverdicted novelty 5.0

    Links WTA training mismatch in GMM-modeled forecasters to uninformative posteriors and introduces post-hoc merging plus one-step EM to yield better-ranked mode probabilities without retraining.

  13. Learning A Unified Risk Map for Autonomous Driving in Partially Observable Environments

    cs.RO 2026-05 unverdicted novelty 5.0

    The paper proposes a unified risk map modeling and learning framework integrated with diffusion-based adversarial scenario generation for risk-aware planning in partially observable autonomous driving, demonstrating i...

  14. Causality-Aware End-to-End Autonomous Driving via Ego-Centric Joint Scene Modeling

    cs.RO 2026-05 unverdicted novelty 5.0

    CaAD adds ego-centric joint-causal modeling and causality-aware policy alignment to end-to-end driving, reporting Driving Score 87.53 and Success Rate 71.81 on Bench2Drive plus PDMS 91.1 on NAVSIM.

  15. Causality-Aware End-to-End Autonomous Driving via Ego-Centric Joint Scene Modeling

    cs.RO 2026-05 unverdicted novelty 5.0

    CaAD adds ego-centric joint-causal modeling and causality-aware policy alignment to end-to-end driving, reporting Driving Score 87.53 and PDMS 91.1 on Bench2Drive and NAVSIM.

  16. SHIELD: Scalable Optimal Control with Certification using Duality and Convexity

    cs.RO 2026-05 unverdicted novelty 5.0

    SHIELD derives safe certificates from Lagrangian duality to reduce decision variables and constraints in convex programs, accelerated by a transformer network, delivering order-of-magnitude speedups in stochastic MPC ...

  17. Recall to Predict: Grounding Motion Forecasting in Interpretable Motion Bank

    cs.CV 2026-05 unverdicted novelty 5.0

    A differentiable motion forecasting model retrieves and refines interpretable trajectory anchors from a contrastively learned motion bank to improve transparency without sacrificing multi-modal accuracy.

  18. Aligning Perception, Reasoning, Modeling and Interaction: A Survey on Physical AI

    cs.AI 2025-10 unverdicted novelty 4.0

    A survey of physical AI that distinguishes theoretical physics reasoning from applied understanding and synthesizes advances in symbolic reasoning, embodied systems, and generative models to advocate for physics-groun...

  19. Transformer-Based Autonomous Driving Models and Deployment-Oriented Compression: A Survey

    cs.LG 2023-04 unverdicted novelty 3.0

    A survey that organizes Transformer-based autonomous driving models by task and architecture while analyzing compression techniques as a system-level deployment concern.