Pith. sign in

REVIEW 3 cited by

SEPT: Towards Efficient Scene Representation Learning for Motion Prediction

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2309.15289 v4 pith:T45B7EQ5 submitted 2023-09-26 cs.CV cs.LG

classification cs.CVcs.LG
keywords forecastingmotionscenesepttrafficagentsargoversecomplex
verification ladder T0 review T1 audit T2 compute T3 formal

Signed reviews

No signed human review yet.

0 comments
read the original abstract

Motion prediction is crucial for autonomous vehicles to operate safely in complex traffic environments. Extracting effective spatiotemporal relationships among traffic elements is key to accurate forecasting. Inspired by the successful practice of pretrained large language models, this paper presents SEPT, a modeling framework that leverages self-supervised learning to develop powerful spatiotemporal understanding for complex traffic scenes. Specifically, our approach involves three masking-reconstruction modeling tasks on scene inputs including agents' trajectories and road network, pretraining the scene encoder to capture kinematics within trajectory, spatial structure of road network, and interactions among roads and agents. The pretrained encoder is then finetuned on the downstream forecasting task. Extensive experiments demonstrate that SEPT, without elaborate architectural design or manual feature engineering, achieves state-of-the-art performance on the Argoverse 1 and Argoverse 2 motion forecasting benchmarks, outperforming previous methods on all main metrics by a large margin.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. GoIRL: Graph-Oriented Inverse Reinforcement Learning for Multimodal Trajectory Prediction

    cs.CV 2025-06 conditional novelty 6.0 of 10

    GoIRL couples maximum entropy inverse reinforcement learning with vectorized lane-graph features to predict multiple future trajectories, reporting competitive benchmark numbers but not the stated state-of-the-art on ...

  2. IMR: Iterative Mode-World Weighted Regression for Multi-Agent Trajectory Prediction

    cs.RO 2026-07 conditional novelty 5.0 of 10

    Mode-world weighted regression and an iterative decoder yield state-of-the-art multi-agent trajectory forecasts on Argoverse 2 by reducing mode collapse while raising ranking and top-1 confidence.

  3. Towards Generalizable Trajectory Prediction Using Dual-Level Representation Learning And Adaptive Prompting

    cs.CV 2025-01 conditional novelty 4.0 of 10

    PerReg+ combines self-distillation, masked reconstruction, register queries, and prompt tuning in a Perceiver-based trajectory predictor, reporting improved accuracy on nuScenes, Argoverse 2, and Waymo.

Pith tools