REVIEW 5 major objections 6 minor 99 references
Leading video generation models have internalized crowd-level pedestrian behavior — density, flow, collision avoidance — but cannot keep individual pedestrians distinct, frequently merging or erasing them.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
A new benchmark (PEDRA) extracts bird's-eye-view pedestrian trajectories from text- and image-generated videos and finds current video models are plausible at crowd level but let pedestrians merge, collide, or disappear.
T0 review reviewed 2026-08-04 challenge →
load-bearing objection A genuinely useful first benchmark for multi-agent pedestrian dynamics in video generation, with a solid I2V arm and an interesting but unvalidated T2V reconstruction pipeline that should not be trusted for absolute numbers yet. the 5 major comments →
PEDRA: Evaluating the Realism of Pedestrian Dynamics in Video Generation
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
Core claim
The paper's central claim is that leading video diffusion models have learned an effective prior for plausible multi-agent pedestrian behavior even though no pedestrian model is built into them. Conditioned on text, they map density and interaction cues onto coherent crowd-level motion: crowded prompts produce larger populations and higher collision rates, directional prompts produce faster walking, and the flow-density curve follows the expected decreasing trend. Conditioned on a start frame, they reproduce the approximate spatial distribution and nearest-neighbor spacing of a ground-truth scene. The same models, however, break at the agent level: pedestrians merge or spontaneously disappea
What carries the argument
The load-bearing piece is a trajectory-extraction pipeline that turns pixel-space detections into metric bird's-eye-view trajectories in synthetic scenes with unknown cameras: an off-the-shelf multi-object tracker finds pedestrians, structure-from-motion estimates per-frame camera pose and scene geometry, a metric-depth estimator supplies real-world scale, a RANSAC alignment fits per-frame scale factors by minimizing a Huber loss between the two depth maps, and an anthropometric check re-scales the mean estimated human height to 1.7 meters whenever it falls outside the 1.4–2.0 meter range. Around this pipeline sits a twelve-metric evaluation protocol spanning trajectory kinematics (velocity,
Load-bearing premise
That the automatic pipeline converting generated pixels into metric bird's-eye-view trajectories — structure-from-motion, metric depth, scale alignment, and a height-based correction — produces world-scale positions accurate enough that the measured speeds, densities, collision rates, and spacing reflect the model's behavior rather than measurement error.
What would settle it
Feed the pipeline a synthetic video rendered with known camera intrinsics, known world-scale human heights, and known pedestrian trajectories. If the recovered bird's-eye trajectories carry mean speed error above a few percent, or if the anthropometric correction fires when the true heights are already correct, then the metric-scale extraction is biased and the quantitative claims about speed and social spacing are not settled.
If this is right
- Text-to-video generation can act as a partial pedestrian simulator: it populates scenes and produces crowd-level motion from a natural-language description, without hand-written rules or parameter tuning.
- The prompt-to-behavior link is measurable: density and interaction words shift population, speed, collision, and flow statistics in consistent directions across five different models.
- Agent permanence is the key bottleneck: any application that counts people or follows individuals across time — evacuation studies, human-robot interaction, exact trajectory prediction — will fail even when the crowd looks right.
- Model design and training data create a real trade-off between single-subject clarity and crowd fidelity; models that filter crowded footage to improve motion clarity show the worst dense-scene behavior.
- The evaluation protocol, and the trajectory-reconstruction method behind it, gives video-generation developers a concrete target: improve agent-level consistency without losing the crowd-level prior.
Where Pith is reading between the lines
- A simple testable extension would make agent integrity a first-class metric: count track terminations that occur while a pedestrian is still in frame, separate from collisions. The paper documents such disappearances qualitatively but does not give them a dedicated score.
- The 1.7 m height re-scaling means the absolute speed numbers inherit the depth estimator's prior. Validating the pipeline on synthetic scenes with known camera intrinsics and known human heights would harden or soften every quantitative comparison in the paper.
- The paper's own Limitations section concedes that the multi-stage extraction pipeline can inject label noise, particularly in metric-scale estimation; the anthropometric correction is a mitigation, not a proof of accuracy.
- The five-second horizon and the observed time-lapse artifacts suggest that longer generations would compound the merging and vanishing failures; a 10–30 second extension would test whether the crowd-level prior degrades with duration.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces PEDRA, a benchmark protocol for evaluating the realism of multi-agent pedestrian dynamics in text-to-video (T2V) and image-to-video (I2V) generation. For I2V, videos are generated from ETH/UCY start frames and compared with a tracker-consistent ground truth through known homographies. For T2V, a 180-prompt suite spanning density and interaction categories is used, and metric-scale bird's-eye-view trajectories are reconstructed from pixel-space using FairMOT, VGGT camera/depth estimates, Depth Pro metric depth, RANSAC scale alignment, and an anthropometric height correction. Twelve metrics cover kinematics, social interaction, and video fidelity. The central claim is that leading video models have learned an effective prior for plausible multi-agent behavior—translating density and interaction prompts into crowd-level motion—but consistently fail at agent-level integrity, with pedestrians merging, disappearing, or becoming untrackable in dense crowds. The paper also provides a large generated-video dataset and a public evaluation protocol.
Significance. If the measurement pipeline is trustworthy, the paper is a useful step toward evaluating video generation models as implicit pedestrian simulators. Its strengths include a structured T2V prompt suite, a large sampled video/track corpus, and a deliberate I2V design that re-processes ground-truth videos with the same MOT pipeline and uses known homographies. The qualitative findings and the I2V comparisons are credible and align with the claimed trade-off between prompt adherence and agent-level consistency. However, the T2V quantitative results are only as strong as the unvalidated metric-scale reconstruction, and two of the fidelity metrics are computed with the same models used to define the trajectories. These issues affect the quantitative support for the central claim, though they do not invalidate the qualitative conclusions or the I2V results.
major comments (5)
- [§3, '3D Reconstruction and Scale Estimation'; Table 3; Appendix C] The absolute T2V metrics (M_vel, M_acc, M_dist, M_flow, M_nn, M_coll) are computed in the reconstructed metric BEV frame. The reconstruction is validated only by the human-height plausibility check and by discarding samples where VGGT and Depth Pro disagree; there is no synthetic ground-truth validation of the scale-alignment pipeline, in contrast to the I2V task. The anthropometric correction constrains the mean height to 1.7 m, but it cannot correct spatially varying depth bias, which would distort per-agent positions and speeds even when the mean height is plausible. Moreover, the depth-consistency discard rates differ across models (WAN 1.78%, HYV 2.33%, CVX 2.78%, LTX 5.0%, OS 3.22%), so Table 3 compares different retained subsets. This is load-bearing for the T2V quantitative claims. The Limitations paragraph acknowledges label noise, but an acknowledgment is not a validation. Plea
- [§3.1, Eq. (20)-(21); Table 1] The video-fidelity metrics are self-referential. M_geo is the mean confidence of VGGT, the same model used to recover camera and depth for the T2V trajectories; M_mot is the FairMOT confidence, and FairMOT is the tracker that determines which tracks enter all metrics. A high M_geo may therefore indicate favorable depth-estimation conditions rather than geometric consistency of the generated video, and M_mot cannot be interpreted independently of the trajectory extraction. Please use an independent reconstruction and tracking model for the fidelity metrics, or at minimum report a correlation analysis and discuss the circularity explicitly.
- [Tables 2 and 3; §4.2] All quantitative results are reported as point estimates with no variance, confidence intervals, or significance tests, despite the multi-stage pipeline and the availability of multiple T2V repetitions and multiple I2V start frames. Statements such as 'WAN demonstrates superior geometric consistency' and 'LTX excels on ZARA2' are based on unquantified differences that may be within pipeline noise. Please provide error bars or bootstrap confidence intervals and, where appropriate, test whether model differences are significant.
- [§3, 'Postprocessing for I2V'; Appendix C] The manuscript is internally inconsistent about the I2V sampling procedure. Section 3 says 'we perform multiple inferences until accumulating at least N_gen = 150 unique tracks or 1500 total detections,' while the I2V description and Appendix C state that one video is generated per start frame, with retries only if no trackable agents are produced. If multiple generations per start frame are in fact used, the 'same start distribution' claim is compromised and the resulting EMD comparisons are biased. Please clarify the exact procedure and, if multiple inferences are used, report the number of generations per start frame and any selection criteria.
- [Appendix A.2, Collision Rate] The collision threshold δ=0.1 m is very small relative to human body width and appears to flag only near-exact overlap of ground-contact points, not actual body collision. Since the collision rate is a headline T2V result, please justify the threshold or report sensitivity to δ over a physically motivated range (e.g., 0.1–0.5 m).
minor comments (6)
- [Table 1] Notation is inconsistent: 'M_DTW_int-div' appears in Table 1 and Eq. (9) in different forms. Please unify.
- [Appendix A.1, Eq. (9)] The expression for Internal Diversity is garbled: the binomial coefficient should be binom{N}{2}. Please fix the math.
- [Appendix A] Typos: 'simulaitons' in the metric introduction and 'compue' in the Velocity paragraph.
- [Figure 10 caption] 'Appenix Fig. 10' in the main text is a typo for 'Appendix Fig. 10'.
- [Table 3] The note that bold values do not indicate desirability is confusing. Use a different visual emphasis (e.g., color or arrows) so that interpretation is not tied to boldness.
- [References] References [96] and [97] appear to be the same VBench-2.0 paper; consider merging them.
Circularity Check
No significant circularity: externally anchored benchmarks and no load-bearing self-citation.
full rationale
The paper's contribution is an evaluation protocol rather than a derived prediction, and its main claims do not reduce by construction to its inputs. The I2V benchmark is externally anchored: trajectories are compared to ETH/UCY ground truth using pre-computed homographies, and the ground truth is re-processed with the same MOT pipeline to avoid tracker-mismatch bias. The T2V benchmark uses VGGT and Depth Pro, both external pretrained models, plus an anthropometric height prior, to recover metric-scale trajectories; the evaluated quantities (velocity, collision rate, density, flow) are not the fitting targets of that scale-estimation procedure. The anthropometric correction imposes mean human height, an external prior, not the behavioral realism scores. The central qualitative conclusions—models can translate density/interaction prompts into crowd-level motion but pedestrians merge or disappear—are supported by visual inspection and by I2V comparisons against ground truth, independent of the T2V scale-estimation details. There are no load-bearing self-citations, no imported uniqueness theorems, and no fitted parameter renamed as a prediction. The Appendix metrics M_mot and M_geo use FairMOT and VGGT confidence as fidelity proxies, and these models also participate in trajectory extraction; the paper explicitly calls them proxies, and they are not used to derive the central behavioral claims. This is a measurement-validity caveat, not a circular derivation. The Limitations section explicitly acknowledges possible label noise in T2V metric scale estimation, confirming that the authors treat this as an error source rather than as an assumed result. Under the stated rules, no circular step can be exhibited with a specific reduction, so the appropriate score is 0.
Axiom & Free-Parameter Ledger
free parameters (8)
- collision distance threshold δ =
0.1 m
- stationary displacement threshold δ_stat =
0.2 m
- flow nearest-neighbor count K =
4
- NN moving threshold ε and search radius =
0.1 m/s, 10 m
- anthropometric mean-height correction =
1.7 m
- per-frame depth scale λ_k =
estimated per frame via RANSAC + Huber loss
- T2V depth-consistency discard thresholds =
≥100 px, ≥30% inliers, residual <10% median metric depth
- I2V sampling target N_gen =
150 unique tracks or 1500 detections
axioms (7)
- domain assumption Pinhole camera model with the bottom-midpoint of each detection bounding box as the ground contact point
- domain assumption VGGT predicts reliable camera intrinsics, extrinsics, and unscaled depth on synthetic generated video
- domain assumption Depth Pro provides accurate metric-scale depth on generated video
- domain assumption FairMOT detects and tracks pedestrians in synthetic video well enough for trajectory metrics
- domain assumption ETH/UCY homographies remain valid for generated I2V clips after filtering for static cameras
- domain assumption Human height and walking-speed priors from literature apply to video-model pedestrians
- domain assumption The fundamental diagram (inverse speed–density relation) is the correct realism target for these scenes
Cite this review
Pith. "Pith review of PEDRA: Evaluating the Realism of Pedestrian Dynamics in Video Generation." pith.science (2026). https://pith.science/paper/AJMUUMUB
@misc{pith2026251020182,
author = {Pith},
title = {Pith review of: PEDRA: Evaluating the Realism of Pedestrian Dynamics in Video Generation},
year = {2026},
howpublished = {\url{https://pith.science/paper/AJMUUMUB}},
note = {Machine review of arXiv:2510.20182}
}
read the original abstract
Pedestrian simulation traditionally relies on expert-tuned, hand-crafted models that limit scalability and generalization. Meanwhile, large-scale video generation models have achieved high visual realism across diverse settings, motivating exploration of their potential as general-purpose world simulators. Existing benchmarks primarily assess single-subject realism rather than scenes with multiple interacting people, leaving the plausibility of multi-agent dynamics in generated videos untested. We propose a rigorous evaluation protocol to benchmark text-to-video (T2V) and image-to-video (I2V) models as implicit simulators of pedestrian dynamics. For I2V, we leverage start frames from established datasets to enable direct comparison with ground truth videos, while for T2V we design a prompt suite covering varied crowd densities and interaction types. A key component is a method to reconstruct 2D bird's-eye view trajectories from pixel-space without known camera parameters. Our analysis shows that leading models exhibit effective priors for plausible multi-agent behavior, though issues such as merging and disappearing pedestrians reveal limits to their physical consistency.
Figures
Reference graph
Works this paper leans on
-
[1]
Crowd management and urban design: New scientific approaches.URBAN DESIGN International, 18(4):282–295, 2013
Kheir Al-Kodmany. Crowd management and urban design: New scientific approaches.URBAN DESIGN International, 18(4):282–295, 2013. 1
2013
-
[2]
Social LSTM: Human Trajectory Prediction in Crowded Spaces
Alexandre Alahi, Kratarth Goel, Vignesh Ramanathan, Alexandre Robicquet, Li Fei-Fei, and Silvio Savarese. Social LSTM: Human Trajectory Prediction in Crowded Spaces. In2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 961–971, 2016. 2
2016
-
[3]
Diffusion for world modeling: Visual details matter in atari
Eloi Alonso, Adam Jelley, Vincent Micheli, Anssi Kan- ervisto, Amos Storkey, Tim Pearce, and Franc ¸ois Fleuret. Diffusion for world modeling: Visual details matter in atari. InAdvances in Neural Information Processing Systems, pages 58757–58791. Curran Associates, Inc., 2024. 3
2024
-
[4]
Improved non-player character (npc) behavior using evolutionary algorithm—a systematic review.Entertainment Computing, 52:100875, 2025
Hendrawan Armanto, Harits Ar Rosyid, Muladi, and Gu- nawan. Improved non-player character (npc) behavior using evolutionary algorithm—a systematic review.Entertainment Computing, 52:100875, 2025. 1
2025
-
[5]
V-jepa 2: Self-supervised video models enable understanding, prediction and planning, 2025
Mido Assran, Adrien Bardes, David Fan, Quentin Garrido, Russell Howes, Mojtaba, Komeili, Matthew Muckley, Am- mar Rizvi, Claire Roberts, Koustuv Sinha, Artem Zholus, Sergio Arnaud, Abha Gejji, Ada Martin, Francois Robert Hogan, Daniel Dugas, Piotr Bojanowski, Vasil Khalidov, Patrick Labatut, Francisco Massa, Marc Szafraniec, Kapil Krishnakumar, Yong Li, X...
2025
-
[6]
Continuous locomotive crowd behavior generation
Inhwan Bae, Junoh Lee, and Hae-Gon Jeon. Continuous locomotive crowd behavior generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2025. 3, 14, 15
2025
-
[7]
Lumiere: A space-time diffusion model for video generation, 2024
Omer Bar-Tal, Hila Chefer, Omer Tov, Charles Her- rmann, Roni Paiss, Shiran Zada, Ariel Ephrat, Junhwa Hur, Guanghui Liu, Amit Raj, Yuanzhen Li, Michael Rubinstein, Tomer Michaeli, Oliver Wang, Deqing Sun, Tali Dekel, and Inbar Mosseri. Lumiere: A space-time diffusion model for video generation, 2024. 3
2024
-
[8]
Beyond ’stampedes’: Towards a new psychology of crowd crush disasters.British Journal of Social Psychology, 63(1):52–69, 2024
Dermot Barr, John Drury, Toby Butler, Sanjeedah Choud- hury, and Fergus Neville. Beyond ’stampedes’: Towards a new psychology of crowd crush disasters.British Journal of Social Psychology, 63(1):52–69, 2024. 1
2024
-
[9]
Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models, 2023
Andreas Blattmann, Robin Rombach, Huan Ling, Tim Dock- horn, Seung Wook Kim, Sanja Fidler, and Karsten Kreis. Align your Latents: High-Resolution Video Synthesis with Latent Diffusion Models, 2023. 3
2023
-
[10]
Richter, and Vladlen Koltun
Aleksei Bochkovskii, Ama ¨el Delaunoy, Hugo Germain, Marcel Santos, Yichao Zhou, Stephan R. Richter, and Vladlen Koltun. Depth pro: Sharp monocular metric depth in less than a second, 2025. 2, 4
2025
-
[11]
Bohannon and A
Richard W. Bohannon and A. Williams Andrews. Normal walking speed: a descriptive meta-analysis.Physiotherapy, 97(3):182–189, 2011. 20
2011
-
[12]
Pyramidal implementation of the affine lucas kanade feature tracker description of the algo- rithm.Intel corporation, 5(1-10):4, 2001
Jean-Yves Bouguet et al. Pyramidal implementation of the affine lucas kanade feature tracker description of the algo- rithm.Intel corporation, 5(1-10):4, 2001. 4
2001
-
[13]
A proposed mathematical model for computer prediction of crowd movements and their associated risks
GE Bradley. A proposed mathematical model for computer prediction of crowd movements and their associated risks. In Proceedings of the International Conference on Engineering for Crowd Safety, pages 303–311. Elsevier Publishing Com- pany London, 1993. 2
1993
-
[14]
Video generation models as world simulators.OpenAI Blog, 1:8, 2024
Tim Brooks, Bill Peebles, Connor Holmes, Will DePue, Yufei Guo, Li Jing, David Schnurr, Joe Taylor, Troy Luh- man, Eric Luhman, et al. Video generation models as world simulators.OpenAI Blog, 1:8, 2024. 1, 3
2024
-
[15]
Chan, Nicolas Heess, Lucy Gonzalez, Simon Osindero, Sherjil Ozair, Scott Reed, Jingwei Zhang, Kon- rad Zolna, Jeff Clune, Nando De Freitas, Satinder Singh, and Tim Rockt¨aschel
Jake Bruce, Michael D Dennis, Ashley Edwards, Jack Parker-Holder, Yuge Shi, Edward Hughes, Matthew Lai, Aditi Mavalankar, Richie Steigerwald, Chris Apps, Yusuf Aytar, Sarah Maria Elisabeth Bechtle, Feryal Behbahani, Stephanie C.Y . Chan, Nicolas Heess, Lucy Gonzalez, Simon Osindero, Sherjil Ozair, Scott Reed, Jingwei Zhang, Kon- rad Zolna, Jeff Clune, Nan...
2024
-
[16]
Caramuta, G
C. Caramuta, G. Collodel, C. Giacomini, C. Gruden, G. Longo, and P. Piccolotto. Survey of detection techniques, mathematical models and simulation software in pedestrian dynamics.Transportation Research Procedia, 25:551–567,
-
[17]
Statistical physics of social dynamics.Rev
Claudio Castellano, Santo Fortunato, and Vittorio Loreto. Statistical physics of social dynamics.Rev. Mod. Phys., 81: 591–646, 2009. 1
2009
-
[18]
Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models
Haoxin Chen, Yong Zhang, Xiaodong Cun, Menghan Xia, Xintao Wang, Chao Weng, and Ying Shan. Videocrafter2: Overcoming data limitations for high-quality video diffu- sion models. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 7310–7320, 2024. 1, 3
2024
-
[19]
Pixart-α: Fast training of dif- fusion transformer for photorealistic text-to-image synthesis,
Junsong Chen, Jincheng Yu, Chongjian Ge, Lewei Yao, Enze Xie, Yue Wu, Zhongdao Wang, James Kwok, Ping Luo, Huchuan Lu, and Zhenguo Li. Pixart-α: Fast training of dif- fusion transformer for photorealistic text-to-image synthesis,
-
[20]
MS-TIP: imputation aware pedestrian trajectory prediction
Pranav Singh Chib, Achintya Nath, Paritosh Kabra, Ishu Gupta, and Pravendra Singh. MS-TIP: imputation aware pedestrian trajectory prediction. InProceedings of the 41st International Conference on Machine Learning. JMLR.org,
-
[21]
Scaling rectified flow trans- formers for high-resolution image synthesis, 2024
Patrick Esser, Sumith Kulal, Andreas Blattmann, Rahim Entezari, Jonas M ¨uller, Harry Saini, Yam Levi, Dominik Lorenz, Axel Sauer, Frederic Boesel, Dustin Podell, Tim Dockhorn, Zion English, Kyle Lacey, Alex Goodwin, Yan- nik Marek, and Robin Rombach. Scaling rectified flow trans- formers for high-resolution image synthesis, 2024. 3 9
2024
-
[22]
Behavioral intention prediction in driving scenes: A survey
Jianwu Fang, Fan Wang, Jianru Xue, and Tat-Seng Chua. Behavioral intention prediction in driving scenes: A survey. IEEE Transactions on Intelligent Transportation Systems, 25 (8):8334–8355, 2024. 1
2024
-
[23]
Crowd-driven mid-scale layout design.ACM Trans
Tian Feng, Lap-Fai Yu, Sai-Kit Yeung, KangKang Yin, and Kun Zhou. Crowd-driven mid-scale layout design.ACM Trans. Graph., 35(4), 2016. 1
2016
-
[24]
Pedestrian planning and design
John J Fruin. Pedestrian planning and design. Technical report, 1971. 20
1971
-
[25]
Seedance 1.0: Exploring the boundaries of video generation models, 2025
Yu Gao, Haoyuan Guo, Tuyen Hoang, Weilin Huang, Lu Jiang, Fangyuan Kong, Huixia Li, Jiashi Li, Liang Li, Xi- aojie Li, Xunsong Li, Yifu Li, Shanchuan Lin, Zhijie Lin, Jiawei Liu, Shu Liu, Xiaonan Nie, Zhiwu Qing, Yuxi Ren, Li Sun, Zhi Tian, Rui Wang, Sen Wang, Guoqiang Wei, Guohong Wu, Jie Wu, Ruiqi Xia, Fei Xiao, Xuefeng Xiao, Jiangqiao Yan, Ceyuan Yang,...
2025
-
[26]
Resolving collisions in dense 3d crowd animations.ACM Trans
Gonzalo Gomez-Nogales, Melania Prieto-Martin, Cris- tian Romero, Marc Comino-Trinidad, Pablo Ramon-Prieto, Anne-H´el`ene Olivier, Ludovic Hoyet, Miguel Otaduy, Julien Pettre, and Dan Casas. Resolving collisions in dense 3d crowd animations.ACM Trans. Graph., 43(5), 2024. 1
2024
-
[27]
Stochastic trajectory prediction via motion indeterminacy diffusion
Tianpei Gu, Guangyi Chen, Junlong Li, Chunze Lin, Yong- ming Rao, Jie Zhou, and Jiwen Lu. Stochastic trajectory prediction via motion indeterminacy diffusion. InProceed- ings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 17113–17122, 2022. 2
2022
-
[28]
Social gan: Socially acceptable tra- jectories with generative adversarial networks
Agrim Gupta, Justin Johnson, Li Fei-Fei, Silvio Savarese, and Alexandre Alahi. Social gan: Socially acceptable tra- jectories with generative adversarial networks. In2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 2255–2264, 2018. 2
2018
-
[29]
Recurrent world models facilitate policy evolution
David Ha and J ¨urgen Schmidhuber. Recurrent world models facilitate policy evolution. InAdvances in Neural Informa- tion Processing Systems. Curran Associates, Inc., 2018. 3
2018
-
[30]
Ltx-video: Realtime video latent diffusion.arXiv preprint arXiv:2501.00103, 2024
Yoav HaCohen, Nisan Chiprut, Benny Brazowski, Daniel Shalem, Dudu Moshe, Eitan Richardson, Eran Levin, Guy Shiran, Nir Zabari, Ori Gordon, Poriya Panet, Sapir Weiss- buch, Victor Kulikov, Yaki Bitterman, Zeev Melumian, and Ofir Bibi. Ltx-video: Realtime video latent diffusion.arXiv preprint arXiv:2501.00103, 2024. 5, 17
Pith/arXiv arXiv 2024
-
[31]
Dream to control: Learning behaviors by la- tent imagination
Danijar Hafner, Timothy Lillicrap, Jimmy Ba, and Moham- mad Norouzi. Dream to control: Learning behaviors by la- tent imagination. InInternational Conference on Learning Representations, 2020. 3
2020
-
[32]
Mastering atari with discrete world mod- els
Danijar Hafner, Timothy P Lillicrap, Mohammad Norouzi, and Jimmy Ba. Mastering atari with discrete world mod- els. InInternational Conference on Learning Representa- tions, 2021. 3
2021
-
[33]
Cambridge university press,
Richard Hartley and Andrew Zisserman.Multiple view ge- ometry in computer vision. Cambridge university press,
-
[34]
CameraCtrl: En- abling Camera Control for Text-to-Video Generation, 2025
Hao He, Yinghao Xu, Yuwei Guo, Gordon Wetzstein, Bo Dai, Hongsheng Li, and Ceyuan Yang. CameraCtrl: En- abling Camera Control for Text-to-Video Generation, 2025. 3, 4
2025
-
[35]
Cameractrl ii: Dynamic scene exploration via camera-controlled video diffusion models, 2025
Hao He, Ceyuan Yang, Shanchuan Lin, Yinghao Xu, Meng Wei, Liangke Gui, Qi Zhao, Gordon Wetzstein, Lu Jiang, and Hongsheng Li. Cameractrl ii: Dynamic scene exploration via camera-controlled video diffusion models, 2025. 4
2025
-
[36]
Videoscore: Building auto- matic metrics to simulate fine-grained human feedback for video generation, 2024
Xuan He, Dongfu Jiang, Ge Zhang, Max Ku, Achint Soni, Sherman Siu, Haonan Chen, Abhranil Chandra, Ziyan Jiang, Aaran Arulraj, Kai Wang, Quy Duc Do, Yuansheng Ni, Bo- han Lyu, Yaswanth Narsupalli, Rongqi Fan, Zhiheng Lyu, Yuchen Lin, and Wenhu Chen. Videoscore: Building auto- matic metrics to simulate fine-grained human feedback for video generation, 2024. 3
2024
-
[37]
Agent-Based Modeling
Dirk Helbing. Agent-Based Modeling. InSocial Self- Organization: Agent-Based Simulations and Experiments to Study Emergent Social Behavior, pages 25–70. Springer, Berlin, Heidelberg, 2012. 1
2012
-
[38]
Social force model for pedestrian dynamics.Phys
Dirk Helbing and P ´eter Moln ´ar. Social force model for pedestrian dynamics.Phys. Rev. E, 51:4282–4286, 1995. 1, 2
1995
-
[39]
Denoising dif- fusion probabilistic models
Jonathan Ho, Ajay Jain, and Pieter Abbeel. Denoising dif- fusion probabilistic models. InAdvances in Neural Infor- mation Processing Systems, pages 6840–6851. Curran Asso- ciates, Inc., 2020. 3
2020
-
[40]
Jonathan Ho, Tim Salimans, Alexey Gritsenko, William Chan, Mohammad Norouzi, and David J. Fleet. Video diffu- sion models, 2022. 3
2022
-
[41]
Autonomous delivery robots: A literature review.IEEE Engineering Management Review, 51(4):77– 89, 2023
Mokter Hossain. Autonomous delivery robots: A literature review.IEEE Engineering Management Review, 51(4):77– 89, 2023. 1
2023
-
[42]
Vbench: Com- prehensive benchmark suite for video generative models
Ziqi Huang, Yinan He, Jiashuo Yu, Fan Zhang, Chenyang Si, Yuming Jiang, Yuanhan Zhang, Tianxing Wu, Qingyang Jin, Nattapol Chanpaisit, Yaohui Wang, Xinyuan Chen, Limin Wang, Dahua Lin, Yu Qiao, and Ziwei Liu. Vbench: Com- prehensive benchmark suite for video generative models. In Proceedings of the IEEE/CVF Conference on Computer Vi- sion and Pattern Reco...
-
[43]
Motion- Diffuser: Controllable Multi-Agent Motion Prediction Us- ing Diffusion
Chiyu ”Max” Jiang, Andre Cornman, Cheolho Park, Ben- jamin Sapp, Yin Zhou, and Dragomir Anguelov. Motion- Diffuser: Controllable Multi-Agent Motion Prediction Us- ing Diffusion. In2023 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 9644–9653, Vancouver, BC, Canada, 2023. IEEE. 1, 2
2023
-
[44]
Fulldit: Multi-task video generative foundation model with full attention, 2025
Xuan Ju, Weicai Ye, Quande Liu, Qiulin Wang, Xintao Wang, Pengfei Wan, Di Zhang, Kun Gai, and Qiang Xu. Fulldit: Multi-task video generative foundation model with full attention, 2025. 1, 3
2025
-
[45]
HunyuanVideo: A Systematic Frame- work For Large Video Generative Models, 2025
Weijie Kong, Qi Tian, Zijian Zhang, Rox Min, Zuozhuo Dai, Jin Zhou, Jiangfeng Xiong, Xin Li, Bo Wu, Jianwei Zhang, Kathrina Wu, Qin Lin, Junkun Yuan, Yanxin Long, Aladdin Wang, Andong Wang, Changlin Li, Duojun Huang, Fang Yang, Hao Tan, Hongmei Wang, Jacob Song, Jiawang Bai, Jianbing Wu, Jinbao Xue, Joey Wang, Kai Wang, Mengyang Liu, Pengyu Li, Shuai Li, ...
2025
-
[46]
Towards social behavior in virtual-agent navigation.Science China Information Sciences, 59(11):112102, 2016
Angelos Kremyzas, Norman Jaklin, and Roland Geraerts. Towards social behavior in virtual-agent navigation.Science China Information Sciences, 59(11):112102, 2016. 1
2016
-
[47]
Crowds by example.Computer Graphics Forum, 26(3):655– 664, 2007
Alon Lerner, Yiorgos Chrysanthou, and Dani Lischinski. Crowds by example.Computer Graphics Forum, 26(3):655– 664, 2007. 2, 3
2007
-
[48]
Generative image dynamics
Zhengqi Li, Richard Tucker, Noah Snavely, and Aleksander Holynski. Generative image dynamics. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 24142–24153, 2024. 3
2024
-
[49]
Evaluation of text-to-video generation models: A dynamics perspective
Mingxiang Liao, Hannan Lu, Xinyu Zhang, Fang Wan, Tianyu Wang, Yuzhong Zhao, Wangmeng Zuo, Qixiang Ye, and Jingdong Wang. Evaluation of text-to-video generation models: A dynamics perspective. InAdvances in Neural In- formation Processing Systems, pages 109790–109816. Cur- ran Associates, Inc., 2024. 3
2024
-
[50]
PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation
Shaowei Liu, Zhongzheng Ren, Saurabh Gupta, and Shen- long Wang. PhysGen: Rigid-Body Physics-Grounded Image-to-Video Generation. InComputer Vision – ECCV 2024, pages 360–378. Springer Nature Switzerland, Cham,
2024
-
[51]
Evalcrafter: Benchmarking and eval- uating large video generation models, 2024
Yaofang Liu, Xiaodong Cun, Xuebo Liu, Xintao Wang, Yong Zhang, Haoxin Chen, Yang Liu, Tieyong Zeng, Raymond Chan, and Ying Shan. Evalcrafter: Benchmarking and eval- uating large video generation models, 2024. 3
2024
-
[52]
An iterative image reg- istration technique with an application to stereo vision
Bruce D Lucas and Takeo Kanade. An iterative image reg- istration technique with an application to stereo vision. In IJCAI’81: 7th international joint conference on Artificial in- telligence, pages 674–679, 1981. 4
1981
-
[53]
Grounding video models to ac- tions through goal conditioned exploration, 2025
Yunhao Luo and Yilun Du. Grounding video models to ac- tions through goal conditioned exploration, 2025. 3
2025
-
[54]
Modeling, evaluation, and scale on artificial pedestrians: A literature review.ACM Comput
Francisco Martinez-Gil, Miguel Lozano, Ignacio Garc ´ıa- Fern´andez, and Fernando Fern ´andez. Modeling, evaluation, and scale on artificial pedestrians: A literature review.ACM Comput. Surv., 50(5), 2017. 1, 3
2017
-
[55]
C. D. Tharindu Mathew, Paulo R. Knob, Soraia Raupp Musse, and Daniel G. Aliaga. Urban walkability design us- ing virtual population simulation.Computer Graphics Fo- rum, 38(1):455–469, 2019. 1
2019
-
[56]
Discovering interaction mechanisms in crowds via deep generative surro- gate experiments.Scientific Reports, 15(1):10385, 2025
Koen Minartz, Fleur Hendriks, Simon Martinus Koop, Alessandro Corbetta, and Vlado Menkovski. Discovering interaction mechanisms in crowds via deep generative surro- gate experiments.Scientific Reports, 15(1):10385, 2025. 5, 14, 19, 20
2025
-
[57]
Social-STGCNN: A Social Spatio- Temporal Graph Convolutional Neural Network for Human Trajectory Prediction
Abduallah Mohamed, Kun Qian, Mohamed Elhoseiny, and Christian Claudel. Social-STGCNN: A Social Spatio- Temporal Graph Convolutional Neural Network for Human Trajectory Prediction. In2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 14412–14420, 2020. 2
2020
-
[58]
Mohler, William B
Betty J. Mohler, William B. Thompson, Sarah H. Creem- Regehr, Herbert L. Pick, and William H. Warren. Visual flow influences gait transition speed and preferred walking speed. Experimental Brain Research, 181(2):221–228, 2007. 20
2007
-
[59]
MotionCraft: Physics- Based Zero-Shot Video Generation
Antonio Montanaro, Luca Savant Aira, Emanuele Aiello, Diego Valsesia, and Enrico Magli. MotionCraft: Physics- Based Zero-Shot Video Generation. InAdvances in Neu- ral Information Processing Systems, pages 123155–123181. Curran Associates, Inc., 2024. 1
2024
-
[60]
Gradeo: Towards human-like evaluation for text- to-video generation via multi-step reasoning, 2025
Zhun Mou, Bin Xia, Zhengchao Huang, Wenming Yang, and Jiaya Jia. Gradeo: Towards human-like evaluation for text- to-video generation via multi-step reasoning, 2025. 3, 17
2025
-
[61]
Mathematical tricks for scalable and ap- pealing crowds in walt disney animation studios’ ”raya and the last dragon”
Nicolas Nghiem. Mathematical tricks for scalable and ap- pealing crowds in walt disney animation studios’ ”raya and the last dragon”. InACM SIGGRAPH 2021 Talks, New York, NY , USA, 2021. Association for Computing Machinery. 1
2021
-
[62]
A Survey of Behavioral Models for Social Robots.Robotics, 8 (3):54, 2019
Olivia Nocentini, Laura Fiorini, Giorgia Acerbi, Alessandra Sorrentino, Gianmaria Mancioppi, and Filippo Cavallo. A Survey of Behavioral Models for Social Robots.Robotics, 8 (3):54, 2019. 1
2019
-
[63]
Cosmos world foundation model platform for physical ai, 2025
NVIDIA, :, Niket Agarwal, Arslan Ali, Maciej Bala, Yogesh Balaji, Erik Barker, Tiffany Cai, Prithvijit Chattopadhyay, Yongxin Chen, Yin Cui, Yifan Ding, Daniel Dworakowski, Jiaojiao Fan, Michele Fenzi, Francesco Ferroni, Sanja Fidler, Dieter Fox, Songwei Ge, Yunhao Ge, Jinwei Gu, Siddharth Gururani, Ethan He, Jiahui Huang, Jacob Huffman, Pooya Jannaty, Ji...
2025
-
[64]
Vasileia Papathanasopoulou, Harris Perakis, Ioanna Spy- ropoulou, and Vassilis Gikas. Pedestrian simulation chal- lenges: Modeling techniques and emerging positioning tech- nologies for its applications.IEEE Transactions on Intelli- gent Transportation Systems, 25(10):12876–12892, 2024. 1
2024
-
[65]
Scalable diffusion models with transformers
William Peebles and Saining Xie. Scalable diffusion models with transformers. InProceedings of the IEEE/CVF inter- national conference on computer vision, pages 4195–4205,
-
[66]
Pellegrini, A
S. Pellegrini, A. Ess, K. Schindler, and L. van Gool. You’ll never walk alone: Modeling social behavior for multi-target tracking. In2009 IEEE 12th International Conference on Computer Vision, pages 261–268, 2009. 2, 3
2009
-
[67]
Open-sora 2.0: Train- ing a commercial-level video generation model in $200k,
Xiangyu Peng, Zangwei Zheng, Chenhui Shen, Tom Young, Xinying Guo, Binluo Wang, Hang Xu, Hongxin Liu, Mingyan Jiang, Wenjun Li, Yuhui Wang, Anbang Ye, Gang Ren, Qianran Ma, Wanying Liang, Xiang Lian, Xiwen Wu, Yuting Zhong, Zhuangyan Li, Chaoyu Gong, Guojun Lei, Leijun Cheng, Limin Zhang, Minghao Li, Ruijie Zhang, Silan Hu, Shijie Huang, Xiaokang Wang, Yu...
-
[68]
Sampson, Shikai Li, Simone Parmeggiani, Steve Fine, Tara Fowler, Vladan Petro- vic, and Yuming Du
Adam Polyak, Amit Zohar, Andrew Brown, Andros Tjandra, Animesh Sinha, Ann Lee, Apoorv Vyas, Bowen Shi, Chih- Yao Ma, Ching-Yao Chuang, David Yan, Dhruv Choudhary, Dingkang Wang, Geet Sethi, Guan Pang, Haoyu Ma, Ishan Misra, Ji Hou, Jialiang Wang, Kiran Jagadeesh, Kunpeng Li, Luxin Zhang, Mannat Singh, Mary Williamson, Matt Le, Matthew Yu, Mitesh Kumar Sin...
2025
-
[69]
Gen3c: 3d-informed world-consistent video generation with precise camera con- trol
Xuanchi Ren, Tianchang Shen, Jiahui Huang, Huan Ling, Yifan Lu, Merlin Nimier-David, Thomas M ¨uller, Alexan- der Keller, Sanja Fidler, and Jun Gao. Gen3c: 3d-informed world-consistent video generation with precise camera con- trol. InProceedings of the IEEE/CVF Conference on Com- puter Vision and Pattern Recognition, 2025. 3, 4
2025
-
[70]
Milacski, Chen Wu, Aayush Prakash, Shingo Takagi, Amaury Aubel, Daeil Kim, Alexandre Bernardino, and Fernando De La Torre
Jose Ribeiro-Gomes, Tianhui Cai, Zolt ´an A. Milacski, Chen Wu, Aayush Prakash, Shingo Takagi, Amaury Aubel, Daeil Kim, Alexandre Bernardino, and Fernando De La Torre. Mo- tionGPT: Human Motion Synthesis with Improved Diversity and Realism via GPT-3 Prompting. In2024 IEEE/CVF Win- ter Conference on Applications of Computer Vision (WACV), pages 5058–5068, ...
2024
-
[71]
High-resolution image synthesis with latent diffusion models
Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Bj ¨orn Ommer. High-resolution image synthesis with latent diffusion models. InProceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10684–10695, 2022. 3
2022
-
[72]
U-net: Convolutional networks for biomedical image segmentation,
Olaf Ronneberger, Philipp Fischer, and Thomas Brox. U-net: Convolutional networks for biomedical image segmentation,
-
[73]
Human height.Our World in Data, 2021
Max Roser, Cameron Appel, and Hannah Ritchie. Human height.Our World in Data, 2021. https://ourworldindata.org/human-height. 5
2021
-
[74]
Rubner, C
Y . Rubner, C. Tomasi, and L.J. Guibas. A metric for dis- tributions with applications to image databases. InSixth International Conference on Computer Vision (IEEE Cat. No.98CH36271), pages 59–66, 1998. 5
1998
-
[75]
Trajectron++: Dynamically-Feasible Trajec- tory Forecasting with Heterogeneous Data
Tim Salzmann, Boris Ivanovic, Punarjay Chakravarty, and Marco Pavone. Trajectron++: Dynamically-Feasible Trajec- tory Forecasting with Heterogeneous Data. InComputer Vi- sion – ECCV 2020, pages 683–700, Cham, 2020. Springer International Publishing. 2
2020
-
[76]
The fundamental diagram of pedestrian move- ment revisited.Journal of Statistical Mechanics: Theory and Experiment, 2005(10):P10002, 2005
Armin Seyfried, Bernhard Steffen, Wolfram Klingsch, and Maik Boltes. The fundamental diagram of pedestrian move- ment revisited.Journal of Statistical Mechanics: Theory and Experiment, 2005(10):P10002, 2005. 5, 20
2005
-
[77]
SGCN: Sparse Graph Convolution Network for Pedestrian Trajectory Pre- diction
Liushuai Shi, Le Wang, Chengjiang Long, Sanping Zhou, Mo Zhou, Zhenxing Niu, and Gang Hua. SGCN: Sparse Graph Convolution Network for Pedestrian Trajectory Pre- diction. In2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8990–8999, 2021. 2
2021
-
[78]
T2v-compbench: A comprehen- sive benchmark for compositional text-to-video generation
Kaiyue Sun, Kaiyi Huang, Xian Liu, Yue Wu, Zihan Xu, Zhenguo Li, and Xihui Liu. T2v-compbench: A comprehen- sive benchmark for compositional text-to-video generation. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 8406–8416,
-
[79]
Qi Sun and Yelda Turkan. A bim-based simulation frame- work for fire safety management and investigation of the critical factors affecting human evacuation performance.Ad- vanced Engineering Informatics, 44:101093, 2020. 1
2020
-
[80]
Holly E Syddall, Leo D Westbury, Cyrus Cooper, and Avan Aihie Sayer. Self-reported walking speed: a useful marker of physical performance among community-dwelling older people?Journal of the American Medical Directors Association, 16(4):323–328, 2015. 20
2015
This paper was first reviewed by deepseek-v4-flash on August 4, 2026.
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.