Pith. sign in

REVIEW 4 major objections 6 minor 94 references

RealTraj: Towards Real-World Pedestrian Trajectory Forecasting

T0 review · 4 major / 6 minor · reviewed 2026-08-12 · deepseek-v4-flash

Pith's one-line read RealTraj claims pedestrian trajectory forecasting can be trained and run on raw detections alone, without person IDs, and still match or beat fully supervised state-of-the-art models.

desk verdict RealTraj is a solid empirical paper showing a detection-only transformer with synthetic pretraining and weakly supervised fine-tuning can hold up against detection and tracking noise; the main open question is whether the closest-detection loss silently tracks the wrong pedestrian in dense crowds. read the letter →

arxiv 2411.17376 v3 pith:YGU723OI submitted 2024-11-26 cs.CV

classification cs.CV
keywords pedestriantrajectoryforecastingdetection-onlyinputself-supervisedpretrainingweaklysupervisedfine-tuningtransformerencoderperceptionerrorrobustnesssyntheticdatapersonidentityannotationcost
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

RealTraj asks whether pedestrian trajectory forecasting can be made practical for real-world deployment by removing two costly assumptions: that perfect tracked trajectories are available, and that every pedestrian must carry a persistent identity label. The paper proposes a two-phase scheme: self-supervised pretraining on synthetic simulator data with three auxiliary tasks (unmasking, denoising, person-identity reconstruction), followed by weakly supervised fine-tuning on real data using only ground-truth detections, where the nearest future detection to each prediction serves as the target and an acceleration regularizer suppresses jitter. The paper reports that this detection-only, ID-free training matches or beats fully supervised state-of-the-art models on JRDB, JTA, and SDD, stays competitive on ETH-UCY, and degrades far less under simulated miss-detections, localization noise, and identity switches. If the claims hold, forecasting models can be trained without expensive person-ID annotations and can run directly on detector output instead of clean tracks.

What carries the argument

The load-bearing design is Det2TrajFormer, a Transformer encoder whose input is a per-frame set of detections (positions only, with no IDs) plus learnable query tokens that read out future positions. Its robustness comes from the removal of identity information from the input stream, making identity-switch errors nonexistent by construction, and from three pretext heads used during synthetic pretraining: an unmasking head that reconstructs masked detections, a denoising head that removes added Gaussian noise, and a person-ID reconstruction head that forces the encoder to associate detections of the same pedestrian across frames. The coupling that makes weak supervision work is the closest-detection loss of Eq. (3)–(4), with the acceleration regularizer of Eq. (5) smoothing the resulting trajectories.

What would settle it

Run the weakly supervised variant on a dense sequence where pedestrians frequently cross paths, and measure how often the predicted trajectories switch to following the wrong person's detections after a crossing, compared with the fully supervised variant; a sizable switch rate would show the closest-detection assumption breaks in crowded scenes.

Watch

Extended reading notes

Core claim

The central claim is that pedestrian trajectory forecasting can be driven entirely by detections, with no person identities ever provided or inferred during training. The proposed model, Det2TrajFormer, is a Transformer encoder that takes an unordered set of past box positions per frame, predicts the future positions of any pedestrian designated as target by a translation of the input frame, and is pretrained on synthetic trajectories with three auxiliary objectives: reconstructing masked detections, denoising corrupted ones, and reconstructing person-identity embeddings so the encoder learns to associate detections across frames. Fine-tuning then uses a weakly supervised loss that matches each predicted future point to the closest ground-truth detection at that timestamp and adds an acceleration-regularization term to prevent oscillation. On the reported benchmarks the fully supervised variant achieves the best ADE/FDE on JRDB, JTA, and SDD, the weakly supervised variant matches fully supervised performance on JRDB and SDD, and under synthetic corruption of 20–80% of inputs the method's error grows far more slowly than that of the leading baselines.

Load-bearing premise

The fine-tuning loss assumes the closest future detection to each predicted point belongs to the same pedestrian; in crowded scenes with crossing paths this can lock predictions onto the wrong person's detections.

Editorial extensions

If this is right

  • Forecasting models can be deployed directly on detector output, eliminating the tracking stage as a prerequisite and the error propagation that comes with it.
  • Person-ID annotation becomes unnecessary for fine-tuning, cutting a major data-preparation cost in building trajectory datasets.
  • Synthetic pretraining with corruption-augmented pretext tasks transfers to real scenes, so real-data collection can be kept small.
  • The acceleration-regularized closest-detection loss is a viable weak supervision signal, giving performance on par with fully supervised training on several benchmarks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The paper leaves open whether the same detection-only principle extends to pedestrians not visible in the last frame; a testable variant would add a no-detection branch to the weak-supervision loss so the model could forecast for occluded agents.
  • The identity-free formulation could be embedded inside a multi-object tracker as a forward motion prior, where its robustness to identity switches might improve association across occlusion gaps, though the paper does not test this.
  • Because the pretext tasks each target one corruption type, one could probe whether matching the corruption ratio used in pretraining to the corruption level expected at deployment changes the robustness curve; the paper ablates the ratio but not this alignment.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 6 minor

Summary. RealTraj is a pedestrian trajectory forecasting framework that takes as input a set of detections without person identities. It consists of Det2TrajFormer, a transformer encoder that processes detection tokens, and two training phases: self-supervised pretraining on synthetic ORCA trajectories with unmasking, denoising, and person-ID reconstruction pretext tasks, followed by weakly-supervised fine-tuning on real ground-truth detections using a nearest-detection regression loss with an acceleration regularizer. The paper reports experiments on JRDB, JTA, ETH-UCY, SDD, and TrajImpute, evaluating robustness to miss-detections, localization errors, and identity switches, few-shot performance, and full-supervision comparisons. The central claims are that the framework reduces ID annotation costs, improves robustness to perception errors, and achieves state-of-the-art or comparable forecasting accuracy on several datasets.

Significance. If validated, the framework addresses an important practical gap: most forecasting models require clean tracked trajectories with consistent IDs, while real perception pipelines produce noisy detections. The paper's strengths include proposing a unified solution to three limitations at once, synthetic pretraining with multiple pretext tasks, and a broad set of experiments covering robustness, few-shot regimes, and ablations. The identity-switch invariance follows naturally from the detection-based input representation, which is a clean architectural design choice. However, the central weakly-supervised objective—the closest-detection target—has a potential failure mode in dense scenes, and the reported experiments do not yet characterize it. The paper would be strengthened substantially by a diagnostic that measures how often the nearest detection is the true target pedestrian.

major comments (4)
  1. [3.4, Eqs. (3)–(5)] The weakly-supervised loss selects the nearest ground-truth detection at each future timestep as the regression target. Because the selection is made independently per timestep without any identity constraint, the target can switch from the target pedestrian to a nearby pedestrian when trajectories cross or in dense crowds. The acceleration regularizer in Eq. (5) only penalizes large second-order differences and cannot prevent a smooth drift that follows different pedestrians. This is load-bearing because it is the mechanism that eliminates person ID annotations. The paper provides no diagnostic for how often d^c_t is the true target pedestrian, nor a density-stratified analysis. Please add such a diagnostic (e.g., the identity-match rate of the nearest detection on the datasets used), and an ablation comparing Eq. (4) with an oracle identity-based target. If the identity-match rate is low in dense scenes, the claims about reducing ID annotation costs need to be revised.
  2. [4.4, Fig. 4] The robustness comparison is not controlled. For miss-detections, the proposed model receives zero-filled detections while the baselines receive linearly interpolated detections; for identity switches, the baselines receive swapped IDs while RealTraj is unaffected by construction because its input has no IDs. As a result, the comparison in Fig. 4 does not isolate model robustness from input preprocessing or architectural assumptions. To support the claim of robustness, the authors should either run all methods on identical corrupted inputs (e.g., the same interpolated or detector-output detections), or explicitly justify why different input treatments are appropriate and add a comparison on realistic detector/tracker outputs.
  3. [Tables 2 and 3] All metrics are reported as single runs without variance. The few-shot experiment in Table 2 randomly selects subsets, so results will vary with the sample; several reported differences are small (e.g., 0.43 vs 0.45 on JRDB at 0.1%). Without standard deviations or repeated seeds, the claimed improvements are not statistically grounded. Please report means and standard deviations over at least three seeds for the few-shot experiments, and ideally for the main comparisons in Table 3.
  4. [Abstract and Table 3] The abstract states that 'the method outperforms state-of-the-art trajectory forecasting methods on multiple datasets.' This holds for the fully-supervised variant on JRDB, JTA, and SDD, but the weakly-supervised variant—the paper's main contribution—does not outperform on ETH-UCY (e.g., minADE20=0.26 vs EqMotion's 0.21; minFDE20=0.43 vs 0.35). The claim should be qualified to distinguish the fully-supervised and weakly-supervised variants, or to state the specific datasets for which the weakly-supervised variant is superior.
minor comments (6)
  1. [Supplementary, Sec. 9] The schedule of the loss weights states that (α, β, γ) changes from (1,0,0) to (0,100,0.1). The value β=100 is surprisingly large and may be a typo; please clarify and, if it is intentional, justify the magnitude.
  2. [Table 2] The w/ Syn. indicators (✓/✗) are embedded within the numeric rows, making the table difficult to parse; please place them in a separate column as described in the caption.
  3. [Sec. 4.3 and Fig. 5(b)] The text says '2K synthetic trajectories' while Fig. 5(b) shows ablation up to 5000 sequences; please clarify the total number of sequences used in the default setting.
  4. [Table 4] The first row (the no-pretraining baseline) is listed with no main task and no pretext tasks; please specify how this model is trained (e.g., from scratch with only the fine-tuning loss) so that the comparison is interpretable.
  5. [Sec. 6] The limitation section mentions only that pedestrians must be detected in the last observed frame; it does not discuss the identity-association ambiguity of the weakly-supervised target in dense crowds, which is a more direct limitation of the proposed fine-tuning scheme.
  6. [Sec. 3.1] The notation X ∈ R^{K T_obs × 2} is inconsistent with X_t ∈ R^{K×2}; consider using a product space or clarifying the reshaping.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the framework is evaluated on held-out ground-truth trajectories and the weakly-supervised loss targets ground-truth detections, not the model's own outputs.

full rationale

RealTraj's central claims are empirical and are validated on held-out test splits of JRDB, JTA, ETH-UCY, SDD, and TrajImpute. The weakly-supervised fine-tuning loss in Eqs. (3)-(4) selects the closest ground-truth future detection d^c_t as the supervision target; this is an external label (albeit a weak one), not the model's own prediction, so the training objective does not reduce to the model's output. The robustness to identity switches is a direct consequence of the architectural choice to consume unlinked detections as input; the paper demonstrates this property by injecting identity swaps and observing that the model is unaffected, which is a design property tested empirically rather than a fitted parameter renamed as a prediction. The self-citations in the paper (e.g., refs. [23] and [24]) are contextual related-work references and are not load-bearing for the main results. No equation in the paper is equivalent to another by construction, and no fitted parameter or validation-set choice is later reported as a predicted quantity. The potential failure mode of the nearest-detection loss attaching to the wrong pedestrian in dense crowds is a correctness and bias concern, not a circularity concern, because the loss still supervises against independent ground-truth detections. Overall, the derivation chain is self-contained with respect to its benchmarks, and no significant circularity is present.

Assumptions & free parameters 4 free parameters · 4 assumptions · 0 invented entities

The framework rests on several hand-chosen hyperparameters and two domain assumptions: that synthetic ORCA data transfers to real crowds, and that the nearest future detection is a valid training target without ID association. These assumptions are reasonable for the presented benchmarks but are not independently verified on raw detector outputs.

free parameters (4)
  • lambda (acceleration regularization weight) = 10
    Tuned on JTA validation (Table 5); set to 10 for all other datasets. This hyperparameter directly controls the smoothness of the weakly-supervised loss.
  • corruption ratio during pretraining = 30%
    Chosen via experiment (Fig. 5a) to maximize downstream ADE on JRDB; corruption is applied as masking/Gaussian noise to past detections.
  • pretext loss weights alpha, beta, gamma = 1,0,0 for first 100 epochs; 0,100,0.1 for final 100 epochs
    Hand-set schedule in supplementary Sec. 9; affects the relative importance of unmasking, denoising, and ID reconstruction.
  • localization noise sigma = 0.5
    Gaussian noise scale used to simulate localization errors during pretraining and robustness evaluation; hand-chosen, not swept.
assumptions (4)
  • domain assumption ORCA synthetic trajectories are representative of real-world pedestrian motion patterns.
    Pretraining relies on the transfer of synthetic data to real benchmarks (Sec. 3.2, Sec. 4.3). If ORCA's collision avoidance and goal-seeking do not match real pedestrian behavior, the pretrained representations may not help.
  • ad hoc to paper Ground-truth future detections are available at fine-tuning time and the closest detection to a predicted point is an adequate surrogate for the target pedestrian's true future position.
    Eqs. (3)-(4) define the weakly-supervised loss using the nearest detection; this bypasses person ID annotation but can match the wrong pedestrian in dense or crossing scenarios.
  • domain assumption Input detections are pre-aligned to the target pedestrian via translation to the origin, and the model can resolve which of K detections in the last frame is the target.
    Sec. 3.1 specifies the target by translating detections; this assumes reliable detection of the target in the last observed frame and that K is correctly known.
  • domain assumption A set-based Transformer input without person identities can learn sufficient association information from positions alone.
    The person ID reconstruction pretext task is designed to recover identity from detections, implying that positions alone carry enough information; if not, the model may fail in highly ambiguous scenes.

how reviews work

0 comments
Cite this review

Pith. "Pith review of RealTraj: Towards Real-World Pedestrian Trajectory Forecasting." pith.science (2026). https://pith.science/paper/YGU723OI

@misc{pith2026241117376,
  author       = {Pith},
  title        = {Pith review of: RealTraj: Towards Real-World Pedestrian Trajectory Forecasting},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/YGU723OI}},
  note         = {Machine review of arXiv:2411.17376}
}
read the original abstract

This paper jointly addresses three key limitations in conventional pedestrian trajectory forecasting: pedestrian perception errors, real-world data collection costs, and person ID annotation costs. We propose a novel framework, RealTraj, that enhances the real-world applicability of trajectory forecasting. Our approach includes two training phases -- self-supervised pretraining on synthetic data and weakly-supervised fine-tuning with limited real-world data -- to minimize data collection efforts. To improve robustness to real-world errors, we focus on both model design and training objectives. Specifically, we present Det2TrajFormer, a trajectory forecasting model that remains invariant to tracking noise by using past detections as inputs. Additionally, we pretrain the model using multiple pretext tasks, which enhance robustness and improve forecasting performance based solely on detection data. Unlike previous trajectory forecasting methods, our approach fine-tunes the model using only ground-truth detections, reducing the need for costly person ID annotations. In the experiments, we comprehensively verify the effectiveness of the proposed method against the limitations, and the method outperforms state-of-the-art trajectory forecasting methods on multiple datasets. The code will be released at https://fujiry0.github.io/RealTraj-project-page.

Figures

Figures reproduced from arXiv: 2411.17376 by the authors.

Figure 1
Figure 1. Our paper addresses the three limitations in the existing [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Our proposed framework consists of two training phases and an inference phase. (1) Self-supervised pretraining on synthetic [PITH_FULL_IMAGE:figures/full_fig_p003_2.png] view at source ↗
Figure 3
Figure 3. Examples of synthetic trajectories generated for self [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (6 more)
Figure 4
Figure 4. Figure 4: Comparison of model robustness against various error types on the JRDB dataset, including detection errors (miss-detections and [PITH_FULL_IMAGE:figures/full_fig_p006_4.png]
Figure 6
Figure 6. Figure 6: The visualization of predicted trajectories on the JRDB [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 5
Figure 5. Figure 5: Effect of corruption ratio during training and number of [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]
Figure 7
Figure 7. Figure 7: Comparison of model robustness against various error types on the JTA dataset, including detection errors (miss-detections and [PITH_FULL_IMAGE:figures/full_fig_p013_7.png]
Figure 8
Figure 8. Figure 8: Comparisons with the current state-of-the-art method [PITH_FULL_IMAGE:figures/full_fig_p013_8.png]
Figure 9
Figure 9. Figure 9: Comparisons with the current state-of-the-art method on [PITH_FULL_IMAGE:figures/full_fig_p013_9.png]

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

94 extracted references · 78 canonical work pages

  1. [1]

    Social LSTM: Human Trajectory Prediction in Crowded Spaces

    Alexandre Alahi, Kratarth Goel, Vignesh Ramanathan, Alexandre Robicquet, Li Fei-Fei, and Silvio Savarese. Social LSTM: Human Trajectory Prediction in Crowded Spaces. In CVPR, 2016. 1, 2, 7

  2. [2]

    A Set of Control Points Conditioned Pedestrian Trajectory Prediction

    Inhwan Bae and Hae-Gon Jeon. A Set of Control Points Conditioned Pedestrian Trajectory Prediction. AAAI, 2023. 1, 6

  3. [3]

    Learning Pedestrian Group Representations for Multi-modal Trajec- tory Prediction

    Inhwan Bae, Jin-Hwi Park, and Hae-Gon Jeon. Learning Pedestrian Group Representations for Multi-modal Trajec- tory Prediction. In ECCV, 2022. 1, 6

  4. [4]

    EigenTrajectory: Low-Rank Descriptors for Multi-Modal Trajectory Forecast- ing

    Inhwan Bae, Jean Oh, and Hae-Gon Jeon. EigenTrajectory: Low-Rank Descriptors for Multi-Modal Trajectory Forecast- ing. In ICCV, 2023. 6, 7

  5. [5]

    Can language beat numerical regression? language-based multimodal tra- jectory prediction

    Inhwan Bae, Junoh Lee, and Hae-Gon Jeon. Can language beat numerical regression? language-based multimodal tra- jectory prediction. In CVPR, 2024. 1

  6. [6]

    Singu- lartrajectory: Universal trajectory predictor using diffusion model

    Inhwan Bae, Young-Jae Park, and Hae-Gon Jeon. Singu- lartrajectory: Universal trajectory predictor using diffusion model. In CVPR, 2024. 1

  7. [7]

    The inD Dataset: A Drone Dataset of Naturalistic Road User Trajectories at German Intersections

    Julian Bock, Robert Krajewski, Tobias Moers, Steffen Runde, Lennart Vater, and Lutz Eckstein. The inD Dataset: A Drone Dataset of Naturalistic Road User Trajectories at German Intersections. In IV, 2020. 1

  8. [8]

    Advdo: Realistic adversarial attacks for trajectory prediction

    Yulong Cao, Chaowei Xiao, Anima Anankuda, Danfei Xu, and Marco Pavone. Advdo: Realistic adversarial attacks for trajectory prediction. In ECCV, 2022. 2

Show all 94 references
  1. [9]

    Traj-MAE: Masked Autoencoders for Trajectory Prediction

    Hao Chen, Jiaze Wang, Kun Shao, Furui Liu, Jianye Hao, Chenyong Guan, Guangyong Chen, and Pheng-Ann Heng. Traj-MAE: Masked Autoencoders for Trajectory Prediction. In ICCV, 2023. 3

  2. [10]

    Mgf: Mixed gaussian flow for diverse trajectory prediction

    Jiahe Chen, Jinkun Cao, Dahua Lin, Kris Kitani, and Jiang- miao Pang. Mgf: Mixed gaussian flow for diverse trajectory prediction. In NeurIPS, 2024. 7

  3. [11]

    Forecast-MAE: Self-supervised pre-training for motion forecasting with masked autoencoders

    Jie Cheng, Xiaodong Mei, and Ming Liu. Forecast-MAE: Self-supervised pre-training for motion forecasting with masked autoencoders. CVPR, 2023. 3, 4

  4. [12]

    Human motion prediction using semi-adaptable neural networks

    Yujiao Cheng, Weiye Zhao, Changliu Liu, and Masayoshi Tomizuka. Human motion prediction using semi-adaptable neural networks. In ACC, 2019. 3

  5. [13]

    Pedestrian Trajec- tory Prediction with Missing Data: Datasets, Imputation, and Benchmarking

    Pranav Singh Chib and Pravendra Singh. Pedestrian Trajec- tory Prediction with Missing Data: Datasets, Imputation, and Benchmarking. In NeurIPS, 2024. 2, 5, 6

  6. [14]

    Multimodal Trajectory Predic- tions for Autonomous Driving using Deep Convolutional Networks

    Henggang Cui, Vladan Radosavljevic, Fang-Chieh Chou, Tsung-Han Lin, Thi Nguyen, Tzu-Kuo Huang, Jeff Schnei- der, and Nemanja Djuric. Multimodal Trajectory Predic- tions for Autonomous Driving using Deep Convolutional Networks. In ICRA, 2019. 1

  7. [15]

    Socially-informed reconstruction for pedestrian trajectory forecasting

    Haleh Damirchi, Ali Etemad, and Michael Greenspan. Socially-informed reconstruction for pedestrian trajectory forecasting. In WACV, 2025. 2

  8. [16]

    BERT: Pre-training of Deep Bidirectional Trans- formers for Language Understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of Deep Bidirectional Trans- formers for Language Understanding. In NAACL, 2019. 3

  9. [17]

    FlowNet: Learn- ing Optical Flow with Convolutional Networks

    Alexey Dosovitskiy, Philipp Fischer, Eddy Ilg, Philip H¨ausser, Caner Hazirbas, Vladimir Golkov, Patrick van der Smagt, Daniel Cremers, and Thomas Brox. FlowNet: Learn- ing Optical Flow with Convolutional Networks. In ICCV,

  10. [18]

    Saits: Self-attention- based imputation for time series

    Wenjie Du, David C ˆot´e, and Yan Liu. Saits: Self-attention- based imputation for time series. Expert Systems with Appli- cations, 2023. 6

  11. [19]

    A review of video surveillance systems

    Omar Elharrouss, Noor Almaadeed, and Somaya Al- Maadeed. A review of video surveillance systems. JVCIR,

  12. [20]

    Learning to Detect and Track Visible and Occluded Body Joints in a Vir- tual World

    Matteo Fabbri, Fabio Lanzi, Simone Calderara, Andrea Palazzi, Roberto Vezzani, and Rita Cucchiara. Learning to Detect and Track Visible and Occluded Body Joints in a Vir- tual World. In ECCV, 2018. 5

  13. [21]

    Mac- Former: Map-Agent Coupled Transformer for Real-Time and Robust Trajectory Prediction

    Chen Feng, Hangning Zhou, Huadong Lin, Zhigang Zhang, Ziyao Xu, Chi Zhang, Boyu Zhou, and Shaojie Shen. Mac- Former: Map-Agent Coupled Transformer for Real-Time and Robust Trajectory Prediction. RAL, 2023. 2, 1

  14. [22]

    Probabilistic Au- tonomous Robot Navigation in Dynamic Environments with Human Motion Prediction

    Amalia Foka and Panos Trahanias. Probabilistic Au- tonomous Robot Navigation in Dynamic Environments with Human Motion Prediction. IJSR, 2010. 1

  15. [23]

    A Two-Block RNN-Based Trajectory Predic- tion From Incomplete Trajectory

    Ryo Fujii, Jayakorn V ongkulbhisal, Ryo Hachiuma, and Hideo Saito. A Two-Block RNN-Based Trajectory Predic- tion From Incomplete Trajectory. IEEE Access, 2021. 2, 1

  16. [24]

    CrowdMAC: Masked Crowd Density Completion for Robust Crowd Den- sity Forecasting

    Ryo Fujii, Ryo Hachiuma, and Hideo Saito. CrowdMAC: Masked Crowd Density Completion for Robust Crowd Den- sity Forecasting. In WACV, 2025. 2

  17. [25]

    Multi- Transmotion: Pre-trained Model for Human Motion Predic- tion

    Yang Gao, Po-Chien Luan, and Alexandre Alahi. Multi- Transmotion: Pre-trained Model for Human Motion Predic- tion. In CoRL, 2024. 3, 4

  18. [26]

    Latent Variable Sequential Set Transformers for Joint Multi-Agent Motion Prediction

    Roger Girgis, Florian Golemo, Felipe Codevilla, Martin Weiss, Jim Aldon D’Souza, Samira Ebrahimi Kahou, Felix Heide, and Christopher Pal. Latent Variable Sequential Set Transformers for Joint Multi-Agent Motion Prediction. In ICLR, 2022. 1, 2, 7

  19. [27]

    Transformer Networks for Trajectory Forecasting

    Francesco Giuliari, Irtiza Hasan, Marco Cristani, and Fabio Galasso. Transformer Networks for Trajectory Forecasting. In ICPR, 2021. 1, 7

  20. [28]

    Stochastic Trajectory Prediction via Motion Indeterminacy Diffusion

    Tianpei Gu, Guangyi Chen, Junlong Li, Chunze Lin, Yong- ming Rao, Jie Zhou, and Jiwen Lu. Stochastic Trajectory Prediction via Motion Indeterminacy Diffusion. In CVPR,

  21. [29]

    Social GAN: Socially Acceptable Tra- jectories with Generative Adversarial Networks

    Agrim Gupta, Justin Johnson, Li Fei-Fei, Silvio Savarese, and Alexandre Alahi. Social GAN: Socially Acceptable Tra- jectories with Generative Adversarial Networks. In CVPR,

  22. [30]

    Geometric trajectory diffusion models

    Jiaqi Han, Minkai Xu, Aaron Lou, Haotian Ye, and Stefano Ermon. Geometric trajectory diffusion models. In NeurIPS,

  23. [31]

    Masked Autoencoders Are Scal- able Vision Learners

    Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ar, and Ross Girshick. Masked Autoencoders Are Scal- able Vision Learners. In CVPR, 2022. 2, 3

  24. [32]

    The Trajectron: Prob- abilistic Multi-Agent Trajectory Modeling With Dynamic Spatiotemporal Graphs

    Boris Ivanovic and Marco Pavone. The Trajectron: Prob- abilistic Multi-Agent Trajectory Modeling With Dynamic Spatiotemporal Graphs. In ICCV, 2019. 2

  25. [33]

    Expand- ing the deployment envelope of behavior prediction via adap- tive meta-learning

    Boris Ivanovic, James Harrison, and Marco Pavone. Expand- ing the deployment envelope of behavior prediction via adap- tive meta-learning. In ICRA, 2023. 3

  26. [34]

    Trajectory prediction: learning to map situations to robot trajectories

    Nikolay Jetchev and Marc Toussaint. Trajectory prediction: learning to map situations to robot trajectories. In ICML,

  27. [35]

    Semi-supervised Semantics-guided Adversarial Training for Robust Trajectory Prediction

    Ruochen Jiao, Xiangguo Liu, Takami Sato, Qi Alfred Chen, and Qi Zhu. Semi-supervised Semantics-guided Adversarial Training for Robust Trajectory Prediction. In ICCV, 2023. 2

  28. [36]

    Pre-training without natural images

    Hirokatsu Kataoka, Kazushige Okayasu, Asato Matsumoto, Eisuke Yamagata, Ryosuke Yamada, Nakamasa Inoue, Akio Nakamura, and Yutaka Satoh. Pre-training without natural images. IJCV, 2022. 3

  29. [37]

    Kingma and Jimmy Ba

    Diederik P. Kingma and Jimmy Ba. Adam: A method for stochastic optimization. In ICLR, 2015. 5

  30. [38]

    Hamid Rezatofighi, and Silvio Savarese

    Vineet Kosaraju, Amir Sadeghian, Roberto Mart ´ın-Mart´ın, Ian Reid, S. Hamid Rezatofighi, and Silvio Savarese. Social- BiGAT: Multimodal Trajectory Forecasting Using Bicycle- GAN and Graph Attention Networks. In NeurIPS, 2019. 2

  31. [39]

    Human Trajectory Forecasting in Crowds: A Deep Learning Per- spective

    Parth Kothari, Sven Kreiss, and Alexandre Alahi. Human Trajectory Forecasting in Crowds: A Deep Learning Per- spective. T-ITS, 2022. 1, 7

  32. [40]

    Learning an Image-Based Motion Context for Multiple People Tracking

    Laura Leal-Taix ´e, Michele Fenzi, Alina Kuznetsova, Bodo Rosenhahn, and Silvio Savarese. Learning an Image-Based Motion Context for Multiple People Tracking. In CVPR,

  33. [41]

    Online multi-agent forecasting with interpretable collaborative graph neural networks

    Maosen Li, Siheng Chen, Yanning Shen, Genjia Liu, Ivor W Tsang, and Ya Zhang. Online multi-agent forecasting with interpretable collaborative graph neural networks. TNNLS,

  34. [42]

    Bcdiff: Bidirectional consistent diffusion for instantaneous trajectory prediction

    Rongqing Li, Changsheng Li, Dongchun Ren, Guangyi Chen, Ye Yuan, and Guoren Wang. Bcdiff: Bidirectional consistent diffusion for instantaneous trajectory prediction. In NeurIPS, 2023. 2, 1

  35. [43]

    LaKD: Length-agnostic Knowl- edge Distillation for Trajectory Prediction with Any Length Observations

    Yuhang Li, Changsheng Li, Ruilin Lv, Rongqing Li, Ye Yuan, and Guoren Wang. LaKD: Length-agnostic Knowl- edge Distillation for Trajectory Prediction with Any Length Observations. In NeurIPS, 2024. 2

  36. [44]

    Zhao, Chenfeng Xu, Chen Tang, Chenran Li, Mingyu Ding, Masayoshi Tomizuka, and Wei Zhan

    Yiheng Li, Seth Z. Zhao, Chenfeng Xu, Chen Tang, Chenran Li, Mingyu Ding, Masayoshi Tomizuka, and Wei Zhan. Pre- training on Synthetic Driving Data for Trajectory Prediction. In IROS, 2024. 3

  37. [45]

    Fast inference and update of probabilistic density estimation on trajectory pre- diction

    Takahiro Maeda and Norimichi Ukita. Fast inference and update of probabilistic density estimation on trajectory pre- diction. In ICCV, 2023. 7

  38. [46]

    It Is Not the Journey But the Destination: Endpoint Conditioned Trajectory Prediction

    Karttikeya Mangalam, Harshayu Girase, Shreyas Agarwal, Kuan-Hui Lee, Ehsan Adeli, Jitendra Malik, and Adrien Gaidon. It Is Not the Journey But the Destination: Endpoint Conditioned Trajectory Prediction. In ECCV, 2020. 2

  39. [47]

    From Goals, Waypoints & Paths to Long Term Human Trajectory Forecasting

    Karttikeya Mangalam, Yang An, Harshayu Girase, and Ji- tendra Malik. From Goals, Waypoints & Paths to Long Term Human Trajectory Forecasting. In ICCV, 2021. 5

  40. [48]

    Leapfrog Diffusion Model for Stochastic Trajectory Prediction

    Weibo Mao, Chenxin Xu, Qi Zhu, Siheng Chen, and Yanfeng Wang. Leapfrog Diffusion Model for Stochastic Trajectory Prediction. In CVPR, 2023. 1, 2

  41. [49]

    Mantra: Memory augmented net- works for multiple trajectory prediction

    Francesco Marchetti, Federico Becattini, Lorenzo Seidenari, and Alberto Del Bimbo. Mantra: Memory augmented net- works for multiple trajectory prediction. In CVPR, 2020. 3

  42. [50]

    Jrdb: A dataset and bench- mark of egocentric robot visual perception of humans in built environments

    Roberto Martin-Martin, Mihir Patel, Hamid Rezatofighi, Abhijeet Shenoi, JunYoung Gwak, Eric Frankel, Amir Sadeghian, and Silvio Savarese. Jrdb: A dataset and bench- mark of egocentric robot visual perception of humans in built environments. TPAMI, 2021. 1, 5

  43. [51]

    Social-STGCNN: A Social Spatio- Temporal Graph Convolutional Neural Network for Human Trajectory Prediction

    Abduallah Mohamed, Kun Qian, Mohamed Elhoseiny, and Christian Claudel. Social-STGCNN: A Social Spatio- Temporal Graph Convolutional Neural Network for Human Trajectory Prediction. In CVPR, 2020. 2

  44. [52]

    How many observations are enough? knowledge distillation for trajec- tory forecasting

    Alessio Monti, Angelo Porrello, Simone Calderara, Pasquale Coscia, Lamberto Ballan, and Rita Cucchiara. How many observations are enough? knowledge distillation for trajec- tory forecasting. In CVPR, 2022. 2, 1

  45. [53]

    DySeT: a Dynamic Masked Self-distillation Approach for Robust Trajectory Prediction

    Amir Rasouli Mozghan Pourkeshavarz, Arielle Zhang. DySeT: a Dynamic Masked Self-distillation Approach for Robust Trajectory Prediction. ECCV, 2024. 3

  46. [54]

    Hager, and Alan L

    Jiteng Mu, Weichao Qiu, Gregory D. Hager, and Alan L. Yuille. Learning from synthetic animals. In CVPR, 2020. 3

  47. [55]

    T4P: Test-Time Training of Trajectory Prediction via Masked Autoencoder and Actor- specific Token Memory

    Daehee Park, Jaeseok Jeong, Sung-Hoon Yoon, Jaewoo Jeong, and Kuk-Jin Yoon. T4P: Test-Time Training of Trajectory Prediction via Masked Autoencoder and Actor- specific Token Memory. In CVPR, 2024. 3

  48. [56]

    Improv- ing Data Association by Joint Modeling of Pedestrian Tra- jectories and Groupings

    Stefano Pellegrini, Andreas Ess, and Luc Van Gool. Improv- ing Data Association by Joint Modeling of Pedestrian Tra- jectories and Groupings. In ECCV, 2010. 1, 5

  49. [57]

    SSL-lanes: Self-supervised learning for motion forecasting in autonomous driving

    Prarthana Bhattacharyya and Chengjie Huang and Krzysztof Czarnecki. SSL-lanes: Self-supervised learning for motion forecasting in autonomous driving. In CoRL, 2022. 3

  50. [58]

    Trace and Pace: Controllable Pedestrian Animation via Guided Trajec- tory Diffusion

    Davis Rempe, Zhengyi Luo, Xue Bin Peng, Ye Yuan, Kris Kitani, Karsten Kreis, Sanja Fidler, and Or Litany. Trace and Pace: Controllable Pedestrian Animation via Guided Trajec- tory Diffusion. In CVPR, 2023. 5, 1

  51. [59]

    Learning Social Etiquette: Human Tra- jectory Understanding In Crowded Scenes

    Alexandre Robicquet, Amir Sadeghian, Alexandre Alahi, and Silvio Savarese. Learning Social Etiquette: Human Tra- jectory Understanding In Crowded Scenes. In ECCV, 2016. 1, 5

  52. [60]

    Kitani, Dariu M

    Andrey Rudenko, Luigi Palmieri, Michael Herman, Kris M. Kitani, Dariu M. Gavrila, and Kai Oliver Arras. Human mo- tion trajectory prediction: a survey. IJRR, 2019. 1

  53. [61]

    Social-Transmotion: Promptable Human Trajectory Prediction

    Saeed Saadatnejad, Yang Gao, Kaouther Messaoud, and Alexandre Alahi. Social-Transmotion: Promptable Human Trajectory Prediction. In ICLR, 2024. 1, 2, 5, 6, 7

  54. [62]

    SoPhie: An Attentive GAN for Predicting Paths Compliant to Social and Physical Constraints

    Amir Sadeghian, Vineet Kosaraju, Ali Sadeghian, Noriaki Hirose, Hamid Rezatofighi, and Silvio Savarese. SoPhie: An Attentive GAN for Predicting Paths Compliant to Social and Physical Constraints. In CVPR, 2019

  55. [63]

    Trajectron++: Dynamically-Feasible Trajec- tory Forecasting with Heterogeneous Data

    Tim Salzmann, Boris Ivanovic, Punarjay Chakravarty, and Marco Pavone. Trajectron++: Dynamically-Feasible Trajec- tory Forecasting with Heterogeneous Data. In ECCV, 2020. 1, 2, 7 10

  56. [64]

    Tra- jectory Unified Transformer for Pedestrian Trajectory Pre- diction

    Liushuai Shi, Le Wang, Sanping Zhou, and Gang Hua. Tra- jectory Unified Transformer for Pedestrian Trajectory Pre- diction. In ICCV, 2023. 1, 6, 7

  57. [65]

    MS-TIP: Imputation aware pedestrian trajectory prediction

    Pranav singh chib, Achintya Nath, Paritosh Kabra, Ishu Gupta, and Pravendra Singh. MS-TIP: Imputation aware pedestrian trajectory prediction. In ICML, 2024. 2, 1

  58. [66]

    Three steps to multimodal trajectory prediction: Modality cluster- ing, classification and synthesis

    Jianhua Sun, Yuxuan Li, Hao-Shu Fang, and Cewu Lu. Three steps to multimodal trajectory prediction: Modality cluster- ing, classification and synthesis. In ICCV, 2021. 3

  59. [67]

    Human trajectory prediction with mo- mentary observation

    Jianhua Sun, Yuxuan Li, Liang Chai, Hao-Shu Fang, Yong- Lu Li, and Cewu Lu. Human trajectory prediction with mo- mentary observation. In CVPR, 2022. 2, 1

  60. [68]

    RAFT: Recurrent All-Pairs Field Transforms for Optical Flow

    Zachary Teed and Jia Deng. RAFT: Recurrent All-Pairs Field Transforms for Optical Flow. In ECCV, 2020. 3

  61. [69]

    Adaptive human trajectory prediction via la- tent corridors

    Neerja Thakkar, Karttikeya Mangalam, Andrea Bajcsy, and Jitendra Malik. Adaptive human trajectory prediction via la- tent corridors. In ECCV, 2024. 3

  62. [70]

    van den Berg, Stephen J

    Jur P. van den Berg, Stephen J. Guy, Ming C Lin, and Dinesh Manocha. Reciprocal n-Body Collision Avoidance. In ISRR,

  63. [71]

    Attention is All you Need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszko- reit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention is All you Need. In NeurIPS, 2017. 4

  64. [72]

    Whose Track Is It Anyway? Improving Robustness to Tracking Errors with Affinity-based Trajectory Prediction

    Xinshuo Weng, Boris Ivanovic, Kris Kitani, and Marco Pavone. Whose Track Is It Anyway? Improving Robustness to Tracking Errors with Affinity-based Trajectory Prediction. In CVPR, 2022. 2, 1

  65. [73]

    MTP: Multi-hypothesis Tracking and Prediction for Reduced Error Propagation

    Xinshuo Weng, Boris Ivanovic, and Marco Pavone. MTP: Multi-hypothesis Tracking and Prediction for Reduced Error Propagation. In IV, 2022. 2, 1

  66. [74]

    Denoising masked autoencoders help ro- bust classification

    QuanLin Wu, Hang Ye, Yuntian Gu, Huishuai Zhang, Liwei Wang, and Di He. Denoising masked autoencoders help ro- bust classification. In ICLR, 2023. 2

  67. [75]

    GroupNet: Multiscale Hypergraph Neural Net- works for Trajectory Prediction With Relational Reasoning

    Chenxin Xu, Maosen Li, Zhenyang Ni, Ya Zhang, and Si- heng Chen. GroupNet: Multiscale Hypergraph Neural Net- works for Trajectory Prediction With Relational Reasoning. In CVPR, 2022. 2

  68. [76]

    PreTraM: Self-Supervised Pre-training via Connecting Tra- jectory and Map

    Chenfeng Xu, Tian Li, Chen Tang, Lingfeng Sun, Kurt Keutzer, Masayoshi Tomizuka, Alireza Fathi, and Wei Zhan. PreTraM: Self-Supervised Pre-training via Connecting Tra- jectory and Map. In ECCV, 2022. 3

  69. [77]

    Remember Intentions: Retrospective-Memory-based Trajec- tory Prediction

    Chenxin Xu, Weibo Mao, Wenjun Zhang, and Siheng Chen. Remember Intentions: Retrospective-Memory-based Trajec- tory Prediction. In CVPR, 2022. 2, 3

  70. [78]

    Tan, Yuhong Tan, Siheng Chen, Xin- chao Wang, and Yanfeng Wang

    Chenxin Xu, Robby T. Tan, Yuhong Tan, Siheng Chen, Xin- chao Wang, and Yanfeng Wang. Auxiliary tasks benefit 3d skeleton-based human motion prediction. In ICCV, 2023. 2

  71. [79]

    Eq- Motion: Equivariant Multi-agent Motion Prediction with In- variant Interaction Reasoning

    Chenxin Xu, Robby T Tan, Yuhong Tan, Siheng Chen, Yu Guang Wang, Xinchao Wang, and Yanfeng Wang. Eq- Motion: Equivariant Multi-agent Motion Prediction with In- variant Interaction Reasoning. In CVPR, 2023. 1, 2, 5, 6, 7

  72. [80]

    Adapting to Length Shift: FlexiLength Network for Trajectory Prediction

    Yi Xu and Yun Fu. Adapting to Length Shift: FlexiLength Network for Trajectory Prediction. In CVPR, 2024. 2, 1

  73. [81]

    Adaptive trajectory prediction via transferable gnn

    Yi Xu, Lichen Wang, Yizhou Wang, and Yun Fu. Adaptive trajectory prediction via transferable gnn. In CVPR, 2022. 3

  74. [82]

    Uncovering the Missing Pattern: Unified Frame- work Towards Trajectory Imputation and Prediction

    Yi Xu, Armin Bazarjani, Hyung-gun Chi, Chiho Choi, and Yun Fu. Uncovering the Missing Pattern: Unified Frame- work Towards Trajectory Imputation and Prediction. In CVPR, 2023. 2, 1

  75. [83]

    Towards Motion Forecasting with Real-World Perception Inputs: Are End-to-End Approaches Competitive? In ICRA, 2024

    Yihong Xu, Lo ¨ıck Chambon, ´Eloi Zablocki, Micka ¨el Chen, Alexandre Alahi, Matthieu Cord, and Patrick P´erez. Towards Motion Forecasting with Real-World Perception Inputs: Are End-to-End Approaches Competitive? In ICRA, 2024. 2

  76. [84]

    Towards Robust Human Trajectory Prediction in Raw Videos

    Rui Yu and Zihan Zhou. Towards Robust Human Trajectory Prediction in Raw Videos. In IROS, 2021. 2, 1

  77. [85]

    Generating Multi-Agent Trajectories using Programmatic Weak Supervision

    Eric Zhan, Stephan Zheng, Yisong Yue, Long Sha, and Patrick Lucey. Generating Multi-Agent Trajectories using Programmatic Weak Supervision. In ICLR, 2019. 1

  78. [86]

    OOSTraj: Out-of-Sight Trajectory Prediction With Vision-Positioning Denoising

    Haichao Zhang, Yi Xu, Hongsheng Lu, Takayuki Shimizu, and Yun Fu. OOSTraj: Out-of-Sight Trajectory Prediction With Vision-Positioning Denoising. In CVPR, 2024. 2, 1

  79. [87]

    TrajPAC: Towards Robustness Verification of Pedestrian Trajectory Prediction Models

    Liang Zhang, Nathaniel Xu, Pengfei Yang, Gaojie Jin, Cheng-Chao Huang, and Lijun Zhang. TrajPAC: Towards Robustness Verification of Pedestrian Trajectory Prediction Models. In ICCV, 2023. 2

  80. [88]

    Towards Trajectory Forecasting From Detection

    Pu Zhang, Lei Bai, Yuning Wang, Jianwu Fang, Jianru Xue, Nanning Zheng, and Wanli Ouyang. Towards Trajectory Forecasting From Detection. TPAMI, 2023. 2, 1

  81. [89]

    On adversarial robustness of tra- jectory prediction for autonomous vehicles

    Qingzhao Zhang, Shengtuo Hu, Jiachen Sun, Qi Alfred Chen, and Z Morley Mao. On adversarial robustness of tra- jectory prediction for autonomous vehicles. In CVPR, 2022. 2

  82. [90]

    Understand- ing collective crowd behaviors: Learning a Mixture model of Dynamic pedestrian-Agents

    Bolei Zhou, Xiaogang Wang, and Xiaoou Tang. Understand- ing collective crowd behaviors: Learning a Mixture model of Dynamic pedestrian-Agents. In CVPR, 2012. 1 11 RealTraj: Towards Real-World Pedestrian Trajectory Forecasting Supplementary Material Table 6. Comparative overvie...

  83. [91]

    Methodological Comparison of Robust Tra- jectory Forecasting against Perception Er- rors To offer a clear and concise comparison between prior works addressing perception errors and our RealTraj ap- proach, we have summarized the key differences in Tab. 6. Unlike previous meth...

  84. [92]

    Up to 40 pedestrians are placed in a 15m × 15m environment with up to 20 static primitive obstacles

    Implementation Details As noted, we use the ORCA crowd simulator [70], imple- mented by [58], to generate synthetic trajectories. Up to 40 pedestrians are placed in a 15m × 15m environment with up to 20 static primitive obstacles. We generate 2000 se- quences for pretraining. ...

  85. [93]

    Comparison with the current state-of-the-art methods on ETH-UCY and SDD for momentary trajectory prediction

    Additional Experimental Results Robustness Evaluation We assess the impact of detection and tracking errors introduced in the JTA dataset, as shown Table 7. Comparison with the current state-of-the-art methods on ETH-UCY and SDD for momentary trajectory prediction. The best re...

  86. [94]

    9, a relatively deep encoder is crucial for optimal performance

    Effect of Number of Encoder Layers As shown in Tab. 9, a relatively deep encoder is crucial for optimal performance. Increasing the number of encoder layers from 3 to 9 results in a 5.3% improvement in ADE. However, adding more layers beyond this point does not yield significa...

Pith tools

Reviewed August 12, 2026 · model on record in the stance chip above.