REVIEW 3 cited by
SEPT: Towards Efficient Scene Representation Learning for Motion Prediction
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
Signed reviews
read the original abstract
Motion prediction is crucial for autonomous vehicles to operate safely in complex traffic environments. Extracting effective spatiotemporal relationships among traffic elements is key to accurate forecasting. Inspired by the successful practice of pretrained large language models, this paper presents SEPT, a modeling framework that leverages self-supervised learning to develop powerful spatiotemporal understanding for complex traffic scenes. Specifically, our approach involves three masking-reconstruction modeling tasks on scene inputs including agents' trajectories and road network, pretraining the scene encoder to capture kinematics within trajectory, spatial structure of road network, and interactions among roads and agents. The pretrained encoder is then finetuned on the downstream forecasting task. Extensive experiments demonstrate that SEPT, without elaborate architectural design or manual feature engineering, achieves state-of-the-art performance on the Argoverse 1 and Argoverse 2 motion forecasting benchmarks, outperforming previous methods on all main metrics by a large margin.
Forward citations
Cited by 3 Pith papers
-
GoIRL: Graph-Oriented Inverse Reinforcement Learning for Multimodal Trajectory Prediction
GoIRL couples maximum entropy inverse reinforcement learning with vectorized lane-graph features to predict multiple future trajectories, reporting competitive benchmark numbers but not the stated state-of-the-art on ...
-
IMR: Iterative Mode-World Weighted Regression for Multi-Agent Trajectory Prediction
Mode-world weighted regression and an iterative decoder yield state-of-the-art multi-agent trajectory forecasts on Argoverse 2 by reducing mode collapse while raising ranking and top-1 confidence.
-
Towards Generalizable Trajectory Prediction Using Dual-Level Representation Learning And Adaptive Prompting
PerReg+ combines self-distillation, masked reconstruction, register queries, and prompt tuning in a Perceiver-based trajectory predictor, reporting improved accuracy on nuScenes, Argoverse 2, and Waymo.
Discussion (0). Continue with ORCID to comment.