REVIEW 3 major objections 6 minor 62 references
SceneDiffuser: Efficient and Controllable Driving Simulation Initialization and Rollout
T0 review · 3 major / 6 minor · reviewed 2026-08-11 · deepseek-v4-flash
Pith's one-line read SceneDiffuser claims that amortized diffusion—running one denoising step per physical simulation step instead of a full denoising loop—cuts closed-loop traffic simulation cost 16-fold and improves realism, while generalized hard…
desk verdict Amortized diffusion is a real contribution with solid internal evidence, but the 'top open-loop' WOSAC claim leans on an unablated validity-mask deviation and a 0.001 margin. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the amortized autoregressive rollout buffer. Instead of re-denoising from pure noise at every physical step, SceneDiffuser warm-starts a future buffer with one one-shot prediction, adds noise under a monotonic schedule $\hat{t}_\tau = \max(0, (\tau - T_{\mathrm{history}})/T_{\mathrm{future}})$, and at each physical step applies one denoising update to the whole buffer, pops the clean first step, and appends a fresh noise sample at the end. Training mixes this monotonic schedule with uniform noise 50/50, so a single model serves both open-loop prediction and amortized closed-loop rollout. The generalized hard constraint mechanism is the second load-bearing piece: a clipping operator applied inside each denoising step that can enforce no-collision, on-road, and feature-range constraints without a differentiable cost.
What would settle it
Re-run the WOSAC closed-loop and one-shot evaluations with a modeled validity mask (or none) and with separate AV/agent rollout steps; if the composite scores drop below the reported 0.673 (closed-loop) and 0.736 (one-shot), or fall behind Trajeglish's 0.735, the claimed leaderboard standing does not survive the protocol change.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that diffusion denoising steps and physical simulation steps can share a time axis: after a one-shot warm-up, the rollout buffer is carried forward and denoised by a single denoising update at each 0.1 s step, with fresh noise appended at the horizon. This amortized autoregressive rollout requires 96 model evaluations over an 80-step rollout (one warm-up plus one per step) versus 1,280 for full autoregressive denoising at 10 Hz, and it scores 0.673 composite on the Waymo Open Sim Agents Challenge metrics versus 0.492 for the full autoregressive baseline at the same replan rate. The same model in one-shot open-loop mode reaches 0.736 composite with the large variant, which the paper reports as top open-loop performance, just ahead of Trajeglish at 0.735, and as the best closed-loop result among diffusion models.
Load-bearing premise
The head-to-head comparisons against the official challenge assume that the paper's stated protocol departure—using the logged validity mask and unifying the AV and agents' rollout step—does not leak future information or inflate scores, and the paper does not ablate that departure.
Editorial extensions
If this is right
- A trained SceneDiffuser model can switch between one-shot open-loop prediction and 10 Hz closed-loop rollout without retraining, since both use the same denoising network and differ only in the noise schedule.
- Closed-loop diffusion simulation becomes practical at full WOSAC scale: amortized rollout needs 96 model evaluations per 8-second scenario instead of 1,280, so the 16x inference reduction translates directly into a much larger feasible replanning rate.
- The realism gap at high replan rates closes: amortized rollout at 10 Hz scores 0.673 composite versus 0.492 for full autoregressive rollout at 10 Hz, so compound-error drift is mitigated rather than merely accepted.
- The open-loop, one-shot mode is itself a competitive motion forecaster (0.736 composite), meaning the same scene prior can be deployed for prediction and simulation.
- Scene editing controls (log perturbation, agent injection, synthetic generation) are available without fine-tuning because they are implemented as inpainting masks and inference-time constraints.
Reading between the lines
- The 16x reduction counts denoising evaluations, not wall-clock time; a runtime benchmark that includes the warm-up and buffer management would show whether the speedup holds in an actual simulator loop.
- If amortized diffusion is as general as the paper suggests, the same single-denoising-step-per-physical-step schedule could be applied to other closed-loop generative world models, such as robotics simulators, where compounding error and inference cost are the same twin obstacles.
- The paper demonstrates hard constraints for scene generation but not for closed-loop rollout; combining GHC with the amortized loop would be a natural next step, since the amortized loop already re-applies a clipping-like operation at every step.
- The comparison to the official leaderboard depends on the logged validity mask and unified rollout step; replacing the logged mask with a learned validity model, which the paper lists as future work, would both remove the protocol deviation and test whether validity modeling is what drives the realism gains.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. SceneDiffuser proposes a unified spatiotemporal diffusion model for traffic simulation that handles both scene initialization (generation, perturbation, agent injection, LLM-constrained generation) and closed-loop rollout. The core methodological contribution is amortized diffusion: denoising steps are aligned with physical simulation steps, so each rollout step requires a single denoising evaluation after a one-shot warm-up, rather than a full autoregressive denoising loop. The paper reports WOSAC realism metrics, ablation and scaling studies, and claims top open-loop performance and the best closed-loop performance among diffusion models. The internal comparison in Table 2 (Amortized AR 0.673 vs. Full AR 0.492 at 10 Hz, with 96 vs. 1280 function evaluations) is the main evidence for the efficiency and closed-loop realism claims.
Significance. If the results hold, the amortized diffusion mechanism is a practical and conceptually useful contribution: it directly addresses the prohibitive inference cost of closed-loop diffusion rollouts and shows that aligning noise levels with physical time reduces compounding error relative to full autoregressive replanning. The unified scene tensor/inpainting formulation across initialization and rollout is also valuable, and the scaling and ablation studies are informative. However, the external WOSAC leaderboard claim rests on a stated protocol deviation that is not ablated, so the headline 'top open-loop performance' cannot currently be taken at face value. The paper is candid in its Limitations section about not beating SOTA autoregressive models, which sharpens the need to resolve the protocol question.
major comments (3)
- [§4.1, Table 4, §A.3] The claimed top open-loop performance (composite 0.736 vs. TRAJEGLISH 0.735) is not supported by a comparable evaluation because of the stated protocol departure: 'we utilize the logged validity mask as input to our transformer and unify the AV and agents' rollout step for simplicity.' The official WOSAC metric in Eqs. (3)-(4) already excludes invalid timesteps via v(i,a,t), but feeding the logged future validity mask as a transformer attention mask tells the model which agents are valid at future timesteps, including agents that enter after the history. This can inflate the composite score by allowing the model to avoid placing likelihood mass on agents that are absent. The 0.001 margin over TRAJEGLISH is well within what such leakage could plausibly contribute. The paper should ablate the one-shot model with and without the logged validity input under the exact official protocol (or with a predicted validity mask), report both numbers, and either keep or remove the 'top open-loop' claim accordingly.
- [§4.1, Table 4, §2.2] The assertion of 'best closed-loop performance among diffusion models' is not established by Table 4. SceneDM and VBD are listed at 0.125 Hz replan and, per §2.2, VBD was evaluated open-loop except on 500 selected scenarios; no other diffusion-based closed-loop rollout is compared. As presented, the comparison reduces to 'our Amortized AR is the best diffusion closed-loop variant we evaluated.' Either provide a direct closed-loop comparison with another diffusion rollout under the same protocol, or rephrase the claim to match what Table 4 actually supports.
- [§5] The Limitations section states 'we do not exceed current SOTA performance for other autoregressive models,' which is in tension with the abstract's 'top open-loop performance.' If the open-loop score is affected by the validity-mask deviation, the abstract overstates the result. The paper should either quantify the effect of the protocol deviation or soften the abstract and contribution statements accordingly.
minor comments (6)
- [§4.1] The word 'degredation' should be 'degradation'.
- [§3.1] The word 'fomulated' should be 'formulated'.
- [Acknowledgments] The phrase 'for detailed for detailed feedback' contains a duplicated fragment and should be corrected.
- [Table 4] Several column headers are malformed or run together (for example 'ADEMINADE'), which makes the per-component results difficult to parse; please reformat the header and explain any bold/color conventions.
- [Fig. 5] The caption states that circle radius is proportional to the number of inference calls, but no scale or numeric labels are provided; adding values would make the efficiency comparison more quantitative.
- [Appendix A.9] The GHC non-collision potential uses a hard-coded 1.5 threshold and an arg-min operation; please state how this optimization is solved at inference time and what its computational overhead is, since GHC is presented as a key controllability mechanism.
Circularity Check
No significant circularity: empirical benchmark claims have independent content; the logged-validity protocol deviation is a correctness caveat, not a circular reduction.
full rationale
SceneDiffuser is an empirical systems paper whose central claims are benchmark evaluations, not derivations from its own outputs. The amortized-diffusion efficiency claim (16x fewer inference steps, Table 2: 1280 vs 96 denoiser evaluations) is a direct accounting of Algorithm 2 versus Algorithm 3, so it is a description of the method rather than a fitted result recycled as a prediction. The realism advantage of Amortized AR over Full AR is measured on the WOSAC benchmark against logged trajectories; no equation in the paper turns a fitted parameter into the reported metric. Self-citations to prior Waymo work, such as MotionDiffuser for the inpainting/scene-tensor formulation and WOSAC for the evaluation protocol, are background methodology and are not load-bearing: the benchmark itself is external, and the stated protocol deviation of using the logged validity mask as a transformer attention mask is a benchmark-validity or correctness concern, not a circularity. The paper itself flags this limitation in its Limitations section: 'We do not explicitly model validity masks and resort to logged validity in this work.' That caveat means the open-loop top ranking should be read with caution, but the ranking does not reduce to the paper's own inputs by construction. No self-citation chain or imported uniqueness theorem forces the main results, so the circularity burden remains low.
Assumptions & free parameters
free parameters (4)
- Feature normalization constants =
scale x,y,z = 1/80; mu_lwhk = [4.5,2.0,1.75,0.5]; sigma_lwhk = [2.5,0.8,0.6,0.5]
- Noise schedule mixing probability =
0.5
- Number of denoising steps =
16
- GHC non-collision potential radius =
1.5 normalized units
assumptions (3)
- standard math The variance-preserving diffusion formalism from Hoogeboom et al. [13] and v-prediction from [37] is valid for the scene tensor representation.
- domain assumption WOMD v1.2.0 is representative of real-world driving, and the WOSAC realism metrics are a meaningful measure of simulation quality.
- domain assumption Providing the logged validity mask as a transformer input at inference, a departure from the official WOSAC setting, does not materially inflate the reported metrics.
Cite this review
Pith. "Pith review of SceneDiffuser: Efficient and Controllable Driving Simulation Initialization and Rollout." pith.science (2026). https://pith.science/paper/FIPIMVYX
@misc{pith2026241212129,
author = {Pith},
title = {Pith review of: SceneDiffuser: Efficient and Controllable Driving Simulation Initialization and Rollout},
year = {2026},
howpublished = {\url{https://pith.science/paper/FIPIMVYX}},
note = {Machine review of arXiv:2412.12129}
}
read the original abstract
Realistic and interactive scene simulation is a key prerequisite for autonomous vehicle (AV) development. In this work, we present SceneDiffuser, a scene-level diffusion prior designed for traffic simulation. It offers a unified framework that addresses two key stages of simulation: scene initialization, which involves generating initial traffic layouts, and scene rollout, which encompasses the closed-loop simulation of agent behaviors. While diffusion models have been proven effective in learning realistic and multimodal agent distributions, several challenges remain, including controllability, maintaining realism in closed-loop simulations, and ensuring inference efficiency. To address these issues, we introduce amortized diffusion for simulation. This novel diffusion denoising paradigm amortizes the computational cost of denoising over future simulation steps, significantly reducing the cost per rollout step (16x less inference steps) while also mitigating closed-loop errors. We further enhance controllability through the introduction of generalized hard constraints, a simple yet effective inference-time constraint mechanism, as well as language-based constrained scene generation via few-shot prompting of a large language model (LLM). Our investigations into model scaling reveal that increased computational resources significantly improve overall simulation realism. We demonstrate the effectiveness of our approach on the Waymo Open Sim Agents Challenge, achieving top open-loop performance and the best closed-loop performance among diffusion models.
Figures
Figures from the paper (8 more)
Reference graph
Works this paper leans on
-
[1]
SimNet: Learning reactive self-driving simulations from real-world observations
Luca Bergamini, Yawei Ye, Oliver Scheel, Long Chen, Chih Hu, Luca Del Pero, Bła˙zej Osi´nski, Hugo Grimmet, and Peter Ondruska. SimNet: Learning reactive self-driving simulations from real-world observations. In ICRA, 2021
work page 2021
-
[2]
Video generation models as world simulators
Tim Brooks, Bill Peebles, Connor Homes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luhman, Eric Luhman, Clarence Wing Yin Ng, Ricky Wang, and Aditya Ramesh. Video generation models as world simulators. 2024. URL https://openai.com/research/ video-generation-models-as-world-simulators
work page 2024
-
[3]
Con- trollable safety-critical closed-loop traffic simulation via guided diffusion, 2023
Wei-Jer Chang, Francesco Pittaluga, Masayoshi Tomizuka, Wei Zhan, and Manmohan Chandraker. Con- trollable safety-critical closed-loop traffic simulation via guided diffusion, 2023
work page 2023
-
[4]
Diffusion policy: Visuomotor policy learning via action diffusion
Cheng Chi, Siyuan Feng, Yilun Du, Zhenjia Xu, Eric Cousineau, Benjamin Burchfiel, and Shuran Song. Diffusion policy: Visuomotor policy learning via action diffusion. In Proceedings of Robotics: Science and Systems (RSS), 2023
2023
-
[5]
Sledge: Synthesizing simulation environments for driving agents with generative models, 2024
Kashyap Chitta, Daniel Dauner, and Andreas Geiger. Sledge: Synthesizing simulation environments for driving agents with generative models, 2024
work page 2024
-
[6]
Dice: Diverse diffusion model with scoring for trajectory prediction, 2023
Younwoo Choi, Ray Coden Mercurius, Soheil Mohamad Alizadeh Shabestary, and Amir Rasouli. Dice: Diverse diffusion model with scoring for trajectory prediction, 2023
work page 2023
-
[7]
Scott Ettinger, Shuyang Cheng, Benjamin Caine, Chenxi Liu, Hang Zhao, Sabeek Pradhan, Yuning Chai, Ben Sapp, Charles R. Qi, Yin Zhou, Zoey Yang, Aurélien Chouard, Pei Sun, Jiquan Ngiam, Vijay Vasudevan, Alexander McCauley, Jonathon Shlens, and Dragomir Anguelov. Large scale interactive motion forecasting for autonomous driving: The waymo open motion datas...
work page 2021
-
[8]
Trafficgen: Learning to generate diverse and realistic traffic scenarios
Lan Feng, Quanyi Li, Zhenghao Peng, Shuhan Tan, and Bolei Zhou. Trafficgen: Learning to generate diverse and realistic traffic scenarios. In ICRA, 2023
work page 2023
Show all 62 references
-
[9]
SceneDM: Scene-level multi-agent trajectory generation with consistent diffusion models, 2023
Zhiming Guo, Xing Gao, Jianlan Zhou, Xinyu Cai, and Botian Shi. SceneDM: Scene-level multi-agent trajectory generation with consistent diffusion models, 2023
2023
-
[10]
Photorealistic video generation with diffusion models, 2023
Agrim Gupta, Lijun Yu, Kihyuk Sohn, Xiuye Gu, Meera Hahn, Li Fei-Fei, Irfan Essa, Lu Jiang, and José Lezama. Photorealistic video generation with diffusion models, 2023
2023
-
[11]
Kingma, Ben Poole, Mohammad Norouzi, David J
Jonathan Ho, William Chan, Chitwan Saharia, Jay Whang, Ruiqi Gao, Alexey Gritsenko, Diederik P. Kingma, Ben Poole, Mohammad Norouzi, David J. Fleet, and Tim Salimans. Imagen video: High definition video generation with diffusion models, 2022
2022
-
[12]
Video diffusion models
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J Fleet. Video diffusion models. In NeurIPS, volume 35, pages 8633–8646, 2022
2022
-
[13]
simple diffusion: End-to-end diffusion for high resolution images
Emiel Hoogeboom, Jonathan Heek, and Tim Salimans. simple diffusion: End-to-end diffusion for high resolution images. In ICML, pages 13213–13232. PMLR, 2023
2023
-
[14]
Versatile scene-consistent traffic scenario generation as optimization with diffusion, 2024
Zhiyu Huang, Zixu Zhang, Ameya Vaidya, Yuxiao Chen, Chen Lv, and Jaime Fernández Fisac. Versatile scene-consistent traffic scenario generation as optimization with diffusion, 2024
2024
-
[15]
Symphony: Learning realistic and diverse agents for autonomous driving simulation
Maximilian Igl, Daewoo Kim, Alex Kuefler, Paul Mougin, Punit Shah, Kyriacos Shiarlis, Dragomir Anguelov, Mark Palatucci, Brandyn White, and Shimon Whiteson. Symphony: Learning realistic and diverse agents for autonomous driving simulation. In ICRA, 2022
2022
-
[16]
Perceiver io: A general architecture for structured inputs & outputs
Andrew Jaegle, Sebastian Borgeaud, Jean-Baptiste Alayrac, Carl Doersch, Catalin Ionescu, David Ding, Skanda Koppula, Daniel Zoran, Andrew Brock, Evan Shelhamer, et al. Perceiver io: A general architecture for structured inputs & outputs. arXiv preprint arXiv:2107.14795, 2021
2021 arXiv
-
[17]
Planning with diffusion for flexible behavior synthesis
Michael Janner, Yilun Du, Joshua Tenenbaum, and Sergey Levine. Planning with diffusion for flexible behavior synthesis. In ICML, 2022
2022
-
[18]
Motiondiffuser: Controllable multi-agent motion prediction using diffusion
Chiyu “Max” Jiang, Andre Cornman, Cheolho Park, Benjamin Sapp, Yin Zhou, and Dragomir Anguelov. Motiondiffuser: Controllable multi-agent motion prediction using diffusion. In CVPR, 2023
2023
-
[19]
Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models
Cheng Lu, Yuhao Zhou, Fan Bao, Jianfei Chen, Chongxuan Li, and Jun Zhu. Dpm-solver++: Fast solver for guided sampling of diffusion probabilistic models. arXiv preprint arXiv:2211.01095, 2022
2022 arXiv
-
[20]
Scenecontrol: Diffusion for controllable traffic scene generation
Jack Lu, Kelvin Wong, Chris Zhang, Simon Suo, and Raquel Urtasun. Scenecontrol: Diffusion for controllable traffic scene generation. In ICRA, 2024. 11
2024
-
[21]
Repaint: Inpainting using denoising diffusion probabilistic models
Andreas Lugmayr, Martin Danelljan, Andres Romero, Fisher Yu, Radu Timofte, and Luc Van Gool. Repaint: Inpainting using denoising diffusion probabilistic models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 11461–11471, June 2022
2022
-
[22]
Unigen: Unified modeling of initial agent states and trajectories for generating autonomous driving scenarios
Reza Mahjourian, Rongbing Mu, Valerii Likhosherstov, Paul Mougin, Xiukun Huang, Joao Messias, and Shimon Whiteson. Unigen: Unified modeling of initial agent states and trajectories for generating autonomous driving scenarios. In ICRA, 2024
2024
-
[23]
The waymo open sim agents challenge
Nico Montali, John Lambert, Paul Mougin, Alex Kuefler, Nick Rhinehart, Michelle Li, Cole Gulino, Tristan Emrich, Zoey Yang, Shimon Whiteson, Brandyn White, and Dragomir Anguelov. The waymo open sim agents challenge. In Advances in Neural Information Processing Systems Track on...
2023
-
[24]
Editable image elements for controllable synthesis
Jiteng Mu, Michaël Gharbi, Richard Zhang, Eli Shechtman, Nuno Vasconcelos, Xiaolong Wang, and Taesung Park. Editable image elements for controllable synthesis. arXiv preprint arXiv:2404.16029, 2024
2024 arXiv
-
[25]
Wayformer: Motion forecasting via simple & efficient attention networks.arXiv preprint arXiv:2207.05844, 2022
Nigamaa Nayakanti, Rami Al-Rfou, Aurick Zhou, Kratarth Goel, Khaled S Refaat, and Benjamin Sapp. Wayformer: Motion forecasting via simple & efficient attention networks.arXiv preprint arXiv:2207.05844, 2022
2022 arXiv
-
[26]
Scene transformer: A unified architecture for predicting future trajectories of multiple agents
Jiquan Ngiam, Vijay Vasudevan, Benjamin Caine, Zhengdong Zhang, Hao-Tien Lewis Chiang, Jeffrey Ling, Rebecca Roelofs, Alex Bewley, Chenxi Liu, Ashish Venugopal, David J Weiss, Benjamin Sapp, Zhifeng Chen, and Jonathon Shlens. Scene transformer: A unified architecture for predi...
2022
-
[27]
A diffusion-model of joint interactive navigation
Matthew Niedoba, Jonathan Lavington, Yunpeng Liu, Vasileios Lioutas, Justice Sefas, Xiaoxuan Liang, Dylan Green, Setareh Dabiri, Berend Zwartsenberg, Adam Scibior, and Frank Wood. A diffusion-model of joint interactive navigation. In NeurIPS, 2023
2023
-
[28]
Lazy diffusion transformer for interactive image editing, 2024
Yotam Nitzan, Zongze Wu, Richard Zhang, Eli Shechtman, Daniel Cohen-Or, Taesung Park, and Michaël Gharbi. Lazy diffusion transformer for interactive image editing, 2024
2024
-
[29]
Imitating human behaviour with diffusion models
Tim Pearce, Tabish Rashid, Anssi Kanervisto, Dave Bignell, Mingfei Sun, Raluca Georgescu, Sergio Val- carcel Macua, Shan Zheng Tan, Ida Momennejad, Katja Hofmann, and Sam Devlin. Imitating human behaviour with diffusion models. In ICLR, 2023
2023
-
[30]
Scalable diffusion models with transformers
William Peebles and Saining Xie. Scalable diffusion models with transformers. InICCV, pages 4195–4205, October 2023
2023
-
[31]
Trajeglish: Learning the language of driving scenarios
Jonah Philion, Xue Bin Peng, and Sanja Fidler. Trajeglish: Learning the language of driving scenarios. In ICLR, 2024
2024
-
[32]
Scenario diffusion: Controllable driving scenario generation with diffusion
Ethan Pronovost, Meghana Reddy Ganesina, Noureldin Hendy, Zeyu Wang, Andres Morales, Kai Wang, and Nick Roy. Scenario diffusion: Controllable driving scenario generation with diffusion. In Advances in Neural Information Processing Systems, 2023
2023
-
[33]
Generating driving scenes with diffusion
Ethan Pronovost, Kai Wang, and Nick Roy. Generating driving scenes with diffusion. In ICRA Workshop on Scalable Autonomous Driving , June 2023
2023
-
[34]
A simple yet effective method for simulating realistic multi-agent behaviors
Cheng Qian, Di Xiu, and Minghao Tian. A simple yet effective method for simulating realistic multi-agent behaviors. Technical report, 2023
2023
-
[35]
Generating useful accident- prone driving scenarios via a learned traffic prior
Davis Rempe, Jonah Philion, Leonidas J Guibas, Sanja Fidler, and Or Litany. Generating useful accident- prone driving scenarios via a learned traffic prior. In CVPR, June 2022
2022
-
[36]
Photorealistic text-to- image diffusion models with deep language understanding
Chitwan Saharia, William Chan, Saurabh Saxena, Lala Li, Jay Whang, Emily L Denton, Kamyar Ghasemipour, Raphael Gontijo Lopes, Burcu Karagol Ayan, Tim Salimans, et al. Photorealistic text-to- image diffusion models with deep language understanding. Advances in neural informatio...
2022
-
[37]
Progressive distillation for fast sampling of diffusion models
Tim Salimans and Jonathan Ho. Progressive distillation for fast sampling of diffusion models. In The Tenth International Conference on Learning Representations, ICLR . OpenReview.net, 2022
2022
-
[38]
Multipath: Multiple probabilistic anchor trajectory hypotheses for behavior prediction
Benjamin Sapp, Yuning Chai, Mayank Bansal, and Dragomir Anguelov. Multipath: Multiple probabilistic anchor trajectory hypotheses for behavior prediction. In Conference on Robot Learning , pages 86–99. PMLR, 2020. 12
2020
-
[39]
Refaat, Rami Al-Rfou, and Benjamin Sapp
Ari Seff, Brian Cera, Dian Chen, Mason Ng, Aurick Zhou, Nigamaa Nayakanti, Khaled S. Refaat, Rami Al-Rfou, and Benjamin Sapp. Motionlm: Multi-agent motion forecasting as language modeling. In Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) , page...
2023
-
[40]
Adafactor: Adaptive learning rates with sublinear memory cost
Noam Shazeer and Mitchell Stern. Adafactor: Adaptive learning rates with sublinear memory cost. In ICML, 2018
2018
-
[41]
Mtr-a: 1st place solution for 2022 waymo open dataset challenge–motion prediction
Shaoshuai Shi, Li Jiang, Dengxin Dai, and Bernt Schiele. Mtr-a: 1st place solution for 2022 waymo open dataset challenge–motion prediction. arXiv preprint arXiv:2209.10033, 2022
2022 arXiv
-
[42]
Mtr++: Multi-agent motion prediction with symmetric scene modeling and guided intention querying
Shaoshuai Shi, Li Jiang, Dengxin Dai, and Bernt Schiele. Mtr++: Multi-agent motion prediction with symmetric scene modeling and guided intention querying. IEEE Transactions on Pattern Analysis and Machine Intelligence, 46(5):3955–3971, 2024
2024
-
[43]
Make-a-video: Text-to-video generation without text-video data
Uriel Singer, Adam Polyak, Thomas Hayes, Xi Yin, Jie An, Songyang Zhang, Qiyuan Hu, Harry Yang, Oron Ashual, Oran Gafni, Devi Parikh, Sonal Gupta, and Yaniv Taigman. Make-a-video: Text-to-video generation without text-video data. In ICLR, 2023
2023
-
[44]
Drivescenegen: Generating diverse and realistic driving scenarios from scratch
Shuo Sun, Zekai Gu, Tianchen Sun, Jiawei Sun, Chengran Yuan, Yuhang Han, Dongen Li, and Marcelo H Ang. Drivescenegen: Generating diverse and realistic driving scenarios from scratch. IEEE Robotics and Automation Letters, 2024
2024
-
[45]
Trafficsim: Learning to simulate realistic multi-agent behaviors
Simon Suo, Sebastian Regalado, Sergio Casas, and Raquel Urtasun. Trafficsim: Learning to simulate realistic multi-agent behaviors. In CVPR, 2021
2021
-
[46]
Scenegen: Learning to generate realistic traffic scenes
Shuhan Tan, Kelvin Wong, Shenlong Wang, Sivabalan Manivasagam, Mengye Ren, and Raquel Urtasun. Scenegen: Learning to generate realistic traffic scenes. In CVPR, June 2021
2021
-
[47]
Language conditioned traffic generation
Shuhan Tan, Boris Ivanovic, Xinshuo Weng, Marco Pavone, and Philipp Krähenbühl. Language conditioned traffic generation. 7th Annual Conference on Robot Learning (CoRL) , 2023
2023
-
[48]
Multipath++: Efficient information fusion and trajectory aggregation for behavior prediction
Balakrishnan Varadarajan, Ahmed Hefny, Avikalp Srivastava, Khaled S Refaat, Nigamaa Nayakanti, Andre Cornman, Kan Chen, Bertrand Douillard, Chi Pang Lam, Dragomir Anguelov, et al. Multipath++: Efficient information fusion and trajectory aggregation for behavior prediction. arX...
2021 arXiv
-
[49]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Ł ukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems, 2017
2017
-
[50]
Nocturne: a scalable driving benchmark for bringing multi-agent learning one step closer to the real world
Eugene Vinitsky, Nathan Lichtlé, Xiaomeng Yang, Brandon Amos, and Jakob Foerster. Nocturne: a scalable driving benchmark for bringing multi-agent learning one step closer to the real world. In NeurIPS Datasets and Benchmarks Track, 2022
2022
-
[51]
Multiverse transformer: 1st place solution for waymo open sim agents challenge 2023
Yu Wang, Tiebiao Zhao, and Fan Yi. Multiverse transformer: 1st place solution for waymo open sim agents challenge 2023. Technical report, Pegasus, 2023
2023
-
[52]
Bits: Bi-level imitation for traffic simulation
Danfei Xu, Yuxiao Chen, Boris Ivanovic, and Marco Pavone. Bits: Bi-level imitation for traffic simulation. In ICRA, 2023
2023
-
[53]
Wcdt: World-centric diffusion transformer for traffic scene generation, 2024
Chen Yang, Aaron Xuxiang Tian, Dong Chen, Tianyu Shi, and Arsalan Heydarian. Wcdt: World-centric diffusion transformer for traffic scene generation, 2024
2024
-
[54]
Trafficbots: Towards world models for autonomous driving simulation and motion prediction
Zhejun Zhang, Alexander Liniger, Dengxin Dai, Fisher Yu, and Luc Van Gool. Trafficbots: Towards world models for autonomous driving simulation and motion prediction. In ICRA, 2023
2023
-
[55]
Tedi: Temporally-entangled diffusion for long-term motion synthesis, 2023
Zihan Zhang, Richard Liu, Kfir Aberman, and Rana Hanocka. Tedi: Temporally-entangled diffusion for long-term motion synthesis, 2023
2023
-
[56]
Language-guided traffic simulation via scene-level diffusion
Ziyuan Zhong, Davis Rempe, Yuxiao Chen, Boris Ivanovic, Yulong Cao, Danfei Xu, Marco Pavone, and Baishakhi Ray. Language-guided traffic simulation via scene-level diffusion. In CoRL, 2023
2023
-
[57]
Guided conditional diffusion for controllable traffic simulation
Ziyuan Zhong, Davis Rempe, Danfei Xu, Yuxiao Chen, Sushant Veer, Tong Che, Baishakhi Ray, and Marco Pavone. Guided conditional diffusion for controllable traffic simulation. In ICRA, 2023. 13 A Appendix / supplemental material A.1 WOSAC Metrics Suppose there are N ≈ 500k scena...
2023
-
[58]
Only t i m e _ s t e p _ i d x values in [0 ,8] are valid for PAST time step
-
[59]
Only t i m e _ s t e p _ i d x values in [0] are valid for CURRENT time step
-
[60]
Only t i m e _ s t e p _ i d x values in [0 ,49] are valid for FUTURE time step
-
[61]
You may only use types POT_CAR , POT_MOTORCYCLIST , P O T _ P E D E S T R I A N to generate these examples
-
[62]
The f ol lo wi ng are 2 examples of a natural language input and the output is a text file that creates the c o r r e s p o n d i n g c o n s t r a i n t
No two agents should overlap each other at the same time step in the same time frame . The f ol lo wi ng are 2 examples of a natural language input and the output is a text file that creates the c o r r e s p o n d i n g c o n s t r a i n t . Example 1 Input : Generate c o n s...
Reviewed August 11, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.