{"id":"ccbde9b8-d75c-40fb-9077-4bc46b6cd0ad","arxiv_id":"2505.18401","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":1.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A brief review of deep learning approaches for crowd behaviour prediction and recognition, with qualitative and quantitative comparisons, concluding that physics-inspired models perform best.","lead":"This chapter reviews recent deep learning methods for crowd behaviour analysis, covering two main tasks: trajectory prediction and behaviour recognition. It organizes the field by neural network architecture and compares representative methods, concluding that physics-inspired deep learning currently leads in accuracy and explainability.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 2's mixed evaluation protocol invalidates the quantitative basis for the claim that physics-inspired deep learning is most accurate.","rationale":"The reader's strongest_claim and weakest_assumption correctly identify the load-bearing issue: the review's central evaluative conclusion is supported primarily by Table 2, and Table 2's comparison is not valid because of the unacknowledged difference between min-over-20 stochastic sampling and deterministic single prediction. This is not a mere presentational flaw; it directly determines whether the 'physics-inspired methods are most accurate' claim can be inferred from the table. The paper itself even flags the problem in the footnote, yet the conclusion proceeds as if the numbers were directly comparable. The same concern is independently reinforced by the fact that, even under the reported numbers, NSP-SFM is not the best on all metrics — NDCPM and IDM beat it on some ADE values — so the conclusion requires a particular weighting and a particular reading of the table. I agree with the reader's conditional verdict: the review is useful as a structured survey and the qualitative discussion of modeling trends has value, but the central comparative claim is not robustly supported. The condition should be a corrected, protocol-controlled comparison before the accuracy conclusion is relied upon. I do not see a need to escalate to rejection, because the concern is about evidence quality rather than an internal mathematical contradiction or a false statement about the existence of the methods. The proposed concrete test — equalizing the sampling protocol and checking whether the ranking survives — would settle whether the concern lands. If the ranking survives, the claim is fine; if it does not, the conclusion must be softened. The reader's verdict already captures this, so no change is needed.","tokens_in":36609,"tokens_out":4421,"duration_ms":39740,"concrete_test":"Re-run the Table 2 comparison under a single evaluation protocol: for every stochastic method, report both the mean and the minimum ADE/FDE over 20 sampled trajectories; for every deterministic method, report the single-prediction error and, where a stochastic variant exists, its min-over-20 error. If NSP-SFM is no longer at or near the top on both ETH/UCY and SDD when protocols are equalized (for example, if IDM or NDCPM wins on ADE or FDE under the same sampling budget), the Section 2.4 conclusion should be weakened from 'strongest' to 'competitive with generative and transformer methods.' If code for all baselines is unavailable, a minimal check is to restrict the ranking to stochastic methods only and compare their min-over-20 and mean-over-20 numbers separately; if the ranking changes, the current claim is not robust.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim in Section 2.4 that physics-inspired deep learning methods 'demonstrate the strongest overall performance, particularly in accuracy' rests on Table 2. Table 2's own footnote states that stochastic methods report the minimum ADE/FDE over 20 sampled trajectories, while deterministic methods (marked with *) report a single prediction. These are not comparable quantities: min-over-samples can only decrease as more samples are drawn, giving stochastic methods a systematic advantage that has nothing to do with model quality. The text in Section 2.4 acknowledges that evaluation metrics vary greatly, then immediately asserts that comparison is possible because the methods 'share common datasets and evaluation metrics' — this is internally inconsistent. Even taking the reported numbers at face value, NSP-SFM is not uniformly best: NDCPM has a lower ETH/UCY ADE (0.15 vs 0.17) and IDM has a lower SDD ADE (6.38 vs 6.52). The conclusion thus depends on weighting FDE and on the unadjusted sampling protocol. In addition, the favorable physics-inspired row is the authors' own NSP-SFM result, so the quantitative evidence is not independent. A corrected comparison is needed before the strong 'strongest overall performance' conclusion can be relied on.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper is a review chapter on recent deep-learning research in crowd behaviour analysis, organized around two core tasks: crowd behaviour prediction and crowd behaviour recognition. For prediction, it surveys traditional statistical machine learning, deep network families (RNN, CNN, GNN, generative, and transformer), and physics-inspired deep learning methods, with detailed case studies of Social-LSTM, NSP-SFM, CrowdMPM, and LDP. For recognition, it covers holistic and individual-based traditional methods as well as CNN, RNN, GNN, and transformer approaches. The chapter concludes that physics-inspired deep learning methods show the strongest overall performance, particularly in accuracy and explainability.","tokens_in":36848,"tokens_out":7131,"duration_ms":54490,"significance":"The review is likely useful as a structured introduction to the field: it covers a broad and current literature, gives explicit equations for representative methods, and offers a consistent taxonomy across both tasks. If the comparative conclusion were properly supported, it would help direct future research toward physics-based priors. However, the central quantitative claim is currently undermined by the protocol mixing in Table 2 and by the non-uniform numbers reported there, so the significance of the conclusion as stated is limited. The descriptive portions and the discussion of future directions (data infrastructure, unsupervised learning, high-density crowds) are the strongest parts of the manuscript.","major_comments":[{"comment":"The footnote to Table 2 states that stochastic methods report the minimum ADE/FDE over 20 sampled trajectories, whereas asterisked deterministic methods report a single deterministic prediction. These quantities are not commensurable: the min over 20 samples improves as more samples are drawn and reflects the best-case prediction, not expected performance. Section 2.4 first notes that \"evaluation metrics and experimental settings vary greatly across different publications\" and then asserts that a numerical comparison is possible because most deep learning methods \"share common datesets and evaluation metrics\"; this is internally contradictory. Since the conclusion that \"physics-inspired deep learning methods demonstrate the strongest overall performance, particularly in accuracy\" rests on Table 2, the ranking must be recomputed under a matched protocol (e.g., best-of-20 for every method, or mean over samples), or the accuracy claim must be downgraded to a qualitative statement.","section":"Table 2; §2.4"},{"comment":"Even accepting the reported values, Table 2 does not establish that physics-inspired methods are uniformly most accurate. NSP-SFM's ETH/UCY ADE (0.17) is worse than NDCPM's (0.15), and its SDD ADE (6.52) is worse than IDM's (6.38). Thus the assertion that physics-inspired deep learning methods \"achieve the highest accuracy\" requires a weighting of ADE against FDE and against dataset-specific performance that is not specified in the text. The conclusion should either name the precise criterion under which NSP-SFM is best (for example, the lowest FDE on both benchmarks) or be softened to \"competitive accuracy with additional explainability benefits.\"","section":"Table 2"},{"comment":"The \"explainability\" part of the claim \"strongest overall performance, particularly in accuracy and explainability\" is not supported by any quantitative or even semi-quantitative evidence in Table 1; the table is a high-level qualitative comparison. The manuscript should define explainability operationally in this context (e.g., the presence of physically interpretable parameters such as forces, masses, and interaction potentials) and should clearly state that the explainability comparison is a judgment, not a measured outcome.","section":"Table 1; §2.4"}],"minor_comments":[{"comment":"The acronym \"NSF-SFM\" appears in the sentence \"Also, since the learn SFM is essentially a simulator, NSF-SFM can simulate more pedestrian behaviours...\"; this should be \"NSP-SFM\" and \"learned SFM.\"","section":"§2.3.1"},{"comment":"There are repeated typos: \"datesets\" and \"common datesets\" should be \"datasets\" and \"common datasets.\"","section":"§2.4 and Table 2 caption"},{"comment":"The citation \"[Velayutham et al.]\" lacks a year and the corresponding reference entry is incomplete; it should be completed with the full bibliographic details.","section":"§4.1 and References"},{"comment":"The sentence \"employing deep learning techniques such as RNNs, CNNs, and GNNs has substantially improved prediction accuracy, reducing ADE and FDE by approximately 20% and 30%, respectively, on ETH/UCY, and by about 50% for both metrics on SDD\" reports summary percentages without showing their computation; either cite the source or show the derivation.","section":"§2.4"},{"comment":"There is a typographical error in \"CrowdMPM He et al. [2025] is the first method of its kind,,\" with a double comma.","section":"§2.3.2"}],"recommendation":"major_revision","confidential_remarks":"The chapter's in-depth methodological sections and the physics-inspired conclusion are substantially based on the authors' own publications (Yue et al. 2022, Yue et al. 2023, He et al. 2025, Yue et al. 2024). In particular, NSP-SFM, the headline accuracy row in Table 2, is the authors' own work. This is not a reason for rejection, but the editor should be aware that the \"strongest overall performance\" claim is partly a self-assessment; the authors should disclose this in the manuscript and ideally include independent replication or a broader set of physics-inspired baselines."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuinely useful survey of deep learning for crowd behaviour analysis, organized by architecture and covering both prediction and recognition. But the stress-test is right that the paper's signature claim—physics-inspired deep learning is the most accurate—is not actually established by the numbers in Table 2.\n\nWhat's good: the taxonomy is sensible (RNN/CNN/GNN/generative/transformer/physics-inspired), the descriptions of representative methods are accurate and accessible, and the recognition half is a fair high-level summary. The discussion of high-density crowds and active-matter approaches is a legitimately interesting frontier. If you are new to the area, this is a fine map.\n\nThe soft spot is exactly where the reader put it. Table 2's footnote says stochastic methods report the minimum ADE/FDE over 20 sampled trajectories while deterministic ones report a single shot. Those numbers are not comparable, and the text on page 23 acknowledges evaluation variance and then immediately claims comparison is possible because methods share datasets/metrics—that's inconsistent. Even putting protocol aside, NSP-SFM is not uniformly best: NDCPM has a lower ETH/UCY ADE (0.15 vs 0.17) and IDM has a lower SDD ADE (6.38 vs 6.52). So the 'strongest overall performance' conclusion depends on weighting FDE more heavily or on an ad hoc aggregate. The fact that the best-looking physics-inspired row is the authors' own NSP-SFM makes it worse. This is a self-citation bias, not fraud, but it weakens the force of the conclusion.\n\nThere is a minor typo ('NSF-SFM' on p.16) and the recognition section has no quantitative comparison at all, so the only load-bearing comparison is the one that is shaky.\n\nWho is this for: newcomers and practitioners wanting a status snapshot. It does not introduce new results; its value is organizational. I would not cite it for the claim that physics-inspired methods win, but I might cite it as an entry-point reference.\n\nRecommendation: send it to peer review, but with a request to fix Table 2—either report a single run for all methods or min-of-20 for all, and show a sensitivity analysis. The authors should also soften the conclusion or present it as a qualitative observation. The review is useful enough to warrant revision rather than rejection.","headline":"Useful review, but the claim that physics-inspired deep learning is most accurate rests on a mixed-protocol table and mostly the authors' own results.","tokens_in":37272,"tokens_out":4611,"would_cite":false,"duration_ms":34756,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":["68T45","68T07"],"pacs":[],"model":"deepseek-v4-flash","headline":"This review argues that embedding physics into neural networks currently gives the strongest overall performance in crowd behaviour prediction, combining accuracy with explainability.","keywords":["crowd behaviour analysis","trajectory prediction","crowd behaviour recognition","physics-inspired deep learning","social force model","neural differential equations","deep learning review","ADE/FDE metrics"],"falsifier":"Re-run the representative methods from Table 2 under one unified evaluation protocol, with the same training and validation splits, the same number of sampled trajectories, and the same random seeds, and check whether NSP-SFM and NDCPM still hold the top ADE/FDE positions. A simpler version: re-evaluate NSP-SFM using the deterministic protocol applied to methods marked with an asterisk, and re-evaluate a deterministic baseline using the min-over-20 protocol; if the ordering between the physics-inspired and pure deep learning families changes, the central claim fails.","tokens_in":36437,"feed_emoji":"🚶","tokens_out":12761,"duration_ms":95240,"temperature":0.7,"pith_summary":"Deep learning has changed how researchers predict and recognise crowd behaviour, and this review chapter maps that landscape by classifying methods according to their network architecture: RNNs, CNNs, GNNs, generative models, transformers, and a newer family that embeds physics into neural networks. Its central conclusion is that the physics-inspired family currently performs best overall in trajectory prediction, delivering the lowest or near-lowest errors on standard benchmarks while also providing something most deep models lack: human-understandable explanations. The authors reach this by combining a qualitative comparison of accuracy, explainability, data requirements, and computational cost with a quantitative table of Average and Final Displacement Error on the ETH/UCY and SDD datasets. If the conclusion holds, it gives safety-critical applications such as autonomous driving and crowd management a concrete reason to favour hybrid physics-neural models over black-box networks.","feed_headline":"Physics-inspired deep learning tops crowd prediction","feed_subtitle":"Embedding physics in neural nets gives top accuracy and explainable forecasts.","key_machinery":"The argument is carried by a two-part comparative structure: a taxonomy that sorts methods by network architecture, and a numerical table that compares representative models on the ETH/UCY and SDD benchmarks using Average Displacement Error and Final Displacement Error. Within that structure, the physics-inspired family's defining mechanism is a neural network whose dynamics are constrained by an explicit physical system: the social force model of Helbing and Molnar embedded as a neural differential equation in NSP-SFM, the material point method with active-matter stress and Toner-Tu active forces in CrowdMPM, and an inverted-pendulum model with learned balance-recovery and interaction forces in LDP. This explicit physical backbone is what the review points to for both the accuracy gains and the explainability advantage.","core_discovery":"The chapter's central claim, stated in the prediction section's conclusion, is that physics-inspired deep learning methods demonstrate the strongest overall performance, particularly in accuracy and explainability. On the quantitative side, two such methods, NSP-SFM and NDCPM, occupy the top of the comparison table, with NSP-SFM reporting an ADE/FDE of 0.17/0.24 on ETH/UCY and 6.52/10.61 on SDD, and NDCPM reporting 0.15/0.33 on ETH/UCY. On the qualitative side, the review rates physics-inspired methods as very high in accuracy and high in explainability, whereas pure deep learning families are rated very high in accuracy but low in explainability. The review also credits the physics component with reducing data requirements, because the physical model supplies structure the network would otherwise have to learn from data.","pith_inferences":["A fair benchmark that standardises the stochastic-versus-deterministic evaluation protocol could reorder the accuracy table; the paper's own footnote leaves this open, so the accuracy lead should be read as provisional until such a benchmark exists.","The physics-inspired advantage may transfer to crowd behaviour recognition, where physics-based features such as entropy, order parameters, and active-Langevin group detection already appear; testing whether physics-informed inductive biases improve recognition accuracy would be a natural extension.","Combining the two dense-crowd directions the review highlights, continuum active-matter models and full-body latent differentiable physics, could yield video-driven risk prediction for crowd crushes, a use case the authors say is currently bottlenecked by data.","Because the review rates physics-inspired methods as needing only medium data, a concrete test is to measure how each architecture family's accuracy degrades as training data shrinks; the physics component should flatten the degradation curve if the claim is right."],"forward_implications":["If the chapter's conclusion is correct, research effort in trajectory prediction should shift toward physics-inspired architectures, since they are the family that combines top accuracy with explainability.","Physical priors reduce the amount of training data needed, making physics-inspired models a better fit for deployment settings where large trajectory datasets do not exist, such as unusual events or new environments.","The physics-inspired principle is portable: continuum and active-matter versions already address dense crowds, and full-body versions address physical perturbations, so the approach is not limited to sparse 2D trajectories.","In safety-critical applications, predictions can be audited: instead of a black-box output, a planner or human operator can inspect the goal-attraction, collision-repulsion, and environment forces that produced a forecast."],"supporting_citations":[{"why":"Introduces NSP-SFM, the physics-inspired neural-social-physics model whose ETH/UCY and SDD results anchor the review's accuracy claim.","marker":"[Yue et al., 2022]"},{"why":"Presents NDCPM, the physics-inspired model that reports the lowest ADE on ETH/UCY in the comparison table.","marker":"[Wang et al., 2024]"},{"why":"Supplies the social force model that NSP-SFM embeds as its physics component, the source of the explainability advantage.","marker":"[Helbing and Molnar, 1995]"},{"why":"Provides the neural ordinary differential equation framework that NSP uses to combine a learnable dynamics model with an explicit physical system.","marker":"[Chen et al., 2018]"},{"why":"Introduces CrowdMPM, the material-point active-matter model that extends physics-inspired deep learning to dense crowds.","marker":"[He et al., 2025]"},{"why":"Introduces LDP, the latent differentiable inverted-pendulum model that extends physics-inspired deep learning to full-body motion under perturbation.","marker":"[Yue et al., 2024]"},{"why":"Supplies the ETH and UCY trajectory datasets that define the ADE/FDE benchmark for the comparison table.","marker":"[Pellegrini et al., 2009, Lerner et al., 2007]"},{"why":"Supplies the SDD dataset whose pixel-based error scores appear in the comparison table.","marker":"[Robicquet et al., 2016]"},{"why":"MLD, the strongest generative-model baseline in the table, whose scores the physics-inspired methods must beat to support the accuracy claim.","marker":"[Wu and Deng, 2024]"},{"why":"PPT, the transformer baseline in the table, representing the best non-physics deep learning family in the comparison.","marker":"[Lin et al., 2024]"}],"fun_headline_variants":["Physics-boosted neural nets lead crowd prediction","Crowd forecasting: physics-inspired AI wins out","Physics + deep learning: top crowd behaviour forecasts","Why physics-infused nets rule crowd analysis","Deep learning with physics tops crowd behaviour prediction"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The ranking assumes that the error scores reported by different papers can be compared directly, even though stochastic methods report the best of 20 random guesses while deterministic methods report a single prediction; if that difference changes the scores, the conclusion that physics-inspired methods are most accurate is unsupported.","fun_headline_variants_meta":{"raw":{"variants":["Physics-boosted neural nets lead crowd prediction","Crowd forecasting: physics-inspired AI wins out","Physics + deep learning: top crowd behaviour forecasts","Why physics-infused nets rule crowd analysis","Deep learning with physics tops crowd behaviour prediction"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000162,"raw_usage":{"total_tokens":1213,"prompt_tokens":896,"completion_tokens":317,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":512,"completion_tokens_details":{"reasoning_tokens":247}},"tokens_in":512,"tokens_out":317,"duration_ms":2976,"temperature":1.0,"reasoning_tokens":247,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T14:31:09.283021+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the representative methods from Table 2 under one unified evaluation protocol, with the same training and validation splits, the same number of sampled trajectories, and the same random seeds, and check whether NSP-SFM and NDCPM still hold the top ADE/FDE positions. A simpler version: re-evaluate NSP-SFM using the deterministic protocol applied to methods marked with an asterisk, and re-evaluate a deterministic baseline using the min-over-20 protocol; if the ordering between the physics-inspired and pure deep learning families changes, the central claim fails.","supporting_citations":[{"cited_title":"Human motion prediction under unexpected perturbation","cited_arxiv_id":null,"evidence_quote":"Introduces LDP, the latent differentiable inverted-pendulum model that extends physics-inspired deep learning to full-body motion under perturbation."},{"cited_title":"Motion latent diffusion for stochastic trajectory prediction","cited_arxiv_id":null,"evidence_quote":"MLD, the strongest generative-model baseline in the table, whose scores the physics-inspired methods must beat to support the accuracy claim."}],"review_version":1}