{"id":"afbb7392-f642-4613-88fd-5646ebabcbd1","arxiv_id":"2507.17317","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"HuNavSim 2.0 is a ROS 2 based simulator that lets users script rich, varied human behaviors with behavior trees and noise-injected crowd models across several robot simulation platforms.","lead":"This paper presents an upgraded open-source simulator for testing how robots navigate around virtual humans. It adds behavior trees, adjustable noise in pedestrian movement, and support for multiple robot simulators, which could help robotics teams develop and benchmark social navigation systems.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Section IV-A's realism gain is asserted, not demonstrated: the parameter sensitivity ranges are undisclosed, the noise sampling scheme is unspecified, and no comparison to real human trajectories is provided.","rationale":"The paper is a tool release; its most defensible value is the modular Behavior Tree framework, the simulator wrappers, and the metrics system, all of which are supported by released code. The weak point is exactly the one the reader identified: the enriched local navigation model is the core novelty that justifies 'realistic' in the title and abstract, and it is neither specified in enough detail to reproduce nor validated. I checked for internal inconsistencies to see whether the concern could be sharpened. Section IV-A references a sensitivity analysis without reporting results; Section V-B contains a two-vs-four metric inconsistency and the total 32, which is a genuine factual error but not load-bearing for the main technical claim. The more serious and precise gap is the underspecified stochastic mechanism: without knowing whether noise is per-agent, per-trial, or per-time-step, the claimed 'variability in navigation trajectories' could be either parameter heterogeneity or unstable dynamics. This does not require rejecting the paper; the tool can be accepted as a community resource if the realism claims are either validated or softened. Since the reader's verdict is already CONDITIONAL and this analysis reinforces that condition, no verdict movement is needed.","tokens_in":7928,"tokens_out":3243,"duration_ms":36149,"concrete_test":"Run a controlled comparison against real pedestrian data. Using the released code, configure the same scenario and population density in deterministic mode and in stochastic mode with default noise distributions, record agent trajectories, and compute trajectory smoothness (e.g., velocity autocorrelation and turning-rate distributions), inter-agent distance statistics, and path predictability. Compare these distributions to the ETH or UCY pedestrian dataset. Also inspect the code to identify the noise sampling frequency and covariance matrix: if sampling occurs every integration step or produces velocity jitter beyond human acceleration bounds, the realism claim in Section IV-A is falsified. If sampling is per-agent only, the 'noise' should be relabeled as parameter heterogeneity and the realism claim should be softened.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central novelty in contribution (ii) is that injecting noise into Social Force Model parameters produces 'variable and more realistic human navigation' (Section IV-A). This claim rests on two unsupported and under-specified pillars. First, the paper states 'We performed a parameter sensitivity analysis to identify feasible ranges that produce small yet realistic variations' but does not report those ranges, the parameters perturbed, or any sensitivity result; the assertion is therefore not independently checkable. Second, the mechanism is ambiguous: it is never stated whether parameter values are sampled once per agent, once per trial, or continuously at each integration step. Under the per-time-step reading, SFM becomes a stochastic differential equation and the added noise may manifest as high-frequency jitter rather than human-like variability; under the per-agent reading, the simulation is simply heterogeneous, not 'noisy' in a way that produces intra-agent behavioral variability. Either way, no quantitative comparison against real pedestrian trajectory data (e.g., ETH/UCY) or against the deterministic baseline is given. The metric-count inconsistency in Section V-B ('adding two new metrics' vs. 'new four metrics' and a total of 32) is a minor factual error that does not affect the main argument, but it reinforces that the evaluation sections were not carefully cross-checked.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents HuNavSim 2.0, an open-source ROS 2-based human navigation simulator for developing and evaluating human-aware robot navigation systems. The new version adds (i) a library of Behavior Tree actions and conditions intended to produce complex and realistic human behaviors, (ii) stochastic noise injection into Social Force Model parameters to introduce trajectory variability, (iii) wrappers for Gazebo Classic, Gazebo Fortress, Nvidia Isaac Sim, and Webots, and (iv) a scenario-creation workflow using RViz 2 panels. The tool also extends the HuNavSim metric suite with additional evaluation metrics. The paper describes these features and gives a qualitative warehouse-worker example, but it contains no quantitative evaluation of the claimed realism improvements or of the tool's behavior-generation capabilities.","tokens_in":8149,"tokens_out":3841,"duration_ms":44206,"significance":"If the claims are substantiated, HuNavSim 2.0 would be a useful community resource: it is open-source, integrates with several widely used robot simulators, provides a Behavior Tree node library, includes a flexible metric system, and ships Docker-based deployment and interactive scenario tools. The multi-simulator wrappers and the RViz2-based scenario editor are practical strengths that lower the entry barrier for human-aware navigation research. However, the two central claims — that parameter noise yields 'variable and more realistic human navigation' and that the Behavior Tree actions yield 'complex and realistic human behaviors' — are presented without any quantitative or comparative evidence. The paper is therefore better evaluated as a system description than as a validated contribution to simulation realism.","major_comments":[{"comment":"The central claim of contribution (ii) — that adding controlled noise to Social Force Model parameters produces 'variable and more realistic human navigation' — is not supported by any reported evidence. The paragraph states that 'We performed a parameter sensitivity analysis to identify feasible ranges that produce small yet realistic variations', but neither the parameters perturbed, the identified ranges, nor the sensitivity results are given. The sampling scheme is also ambiguous: it is not stated whether parameter values are drawn once per agent, once per simulation run, or at every integration step. These regimes have very different consequences: per-agent sampling yields a heterogeneous but deterministic population, whereas per-step sampling turns the SFM into a stochastic differential equation that may produce high-frequency jitter rather than smooth human-like variability. Finally, there is no comparison against real pedestrian trajectory data (e.g., ETH/UCY) or against the deterministic SFM baseline. The realism claim is therefore not independently checkable and needs to be backed by quantitative evidence, including a specification of the noise model and an analysis of the resulting trajectory distributions.","section":"Section IV-A"},{"comment":"The abstract and conclusions state that the Behavior Tree actions and conditions 'compound complex and realistic human behaviors', but the only supporting evidence is a narrative description of a warehouse scenario involving two workers. No evaluation is provided to show that the BT-enabled behaviors are realistic or that they go beyond what the SFM alone can generate. Since 'realistic' is a stated contribution in both the introduction and the conclusion, the paper should include at least a qualitative validation (e.g., a user study, comparison with real human behavior data, or an analysis of behavior variability) to justify this claim. Without such evidence, the BT contribution rests on assertion rather than demonstration.","section":"Sections IV-B and VI"}],"minor_comments":[{"comment":"There is a numerical inconsistency in the metric count: the text says 'adding two new interesting metrics', then refers to 'The new four metrics', and then states 'That gives a total of 32 metrics' when the base is said to be 28. Since 28 + 4 = 32, the 'two' is likely a typo, but the contradiction should be corrected for clarity.","section":"Section V-B"},{"comment":"The title in the full text reads 'A Enhanced Human Navigation Simulator' and should be 'An Enhanced Human Navigation Simulator'.","section":"Title and header"},{"comment":"Several typesetting errors occur: 'InHuNavSim' and 'HuNavSim2.0' are missing spaces in the introduction and conclusion; 'Warehoure' in the Fig. 4 caption should be 'Warehouse'; 'botton' should be 'bottom' in the same caption.","section":"Throughout"},{"comment":"The paper claims the metric set is 'the most comprehensive collection of metrics for human-aware navigation to date' but does not compare the metric list with those of SocNavBench, CrowdBot, or other recent benchmarks beyond citing them. A succinct comparison table would make this claim verifiable.","section":"Section V-B"},{"comment":"The example Behavior Tree for Worker 2 is described in detail in the text, but the corresponding figure (Fig. 3) is not discussed step-by-step; referring to node names in the figure would help readers connect the narrative to the tree structure.","section":"Section IV-B"}],"recommendation":"major_revision","confidential_remarks":"This is a tool paper in which the core contributions are primarily engineering. The main risk is that the paper overclaims realism without supporting data. The missing sensitivity analysis and the ambiguity in the noise sampling mechanism are fixable with a moderate amount of additional experimentation. I would encourage the editor to ask for a revised version that includes quantitative evidence for the noise-based variability claim and at least a qualitative evaluation of the Behavior Tree behaviors, while noting that the metric-count inconsistency suggests the evaluation sections need careful re-reading. The paper's fit with a robotics venue is otherwise reasonable, given its open-source release and multi-simulator support."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"HuNavSim 2.0 is a genuine engineering contribution: it extends the authors' own open-source simulator with a proper Behavior Tree node library, SFM parameter noise, wrappers for four robot simulators, and an RViz2-based scenario editor. If you work on social navigation benchmarking, this is probably the most flexible tool of its kind right now. The code and docs are public, and the metric set (32 total) is genuinely broader than what CrowdBot, SEAN, or SocNavBench offer. I'd use it.\n\nWhat's new here is not the component ideas—SFM with noise and BT control are established—but the packaging and the breadth of integration. That is worth something, and the paper is honest about building on [1]. The warehouse BT example is a nice concrete demonstration.\n\nThe soft spots are exactly where the reader put them. The central claim in IV-A—that Gaussian noise on SFM parameters yields 'variable and more realistic human navigation'—is asserted, not shown. The claimed sensitivity analysis is mentioned but never reported: no ranges, no parameters, no method, and no comparison against a deterministic baseline or real pedestrian data (ETH/UCY would be the obvious choice). Worse, the sampling scheme is ambiguous: per-agent, per-trial, or per time-step? That ambiguity matters because the per-step reading turns SFM into an SDE with possible jitter rather than behavioral variability. This is a load-bearing weakness for the realism claim, but it is not a flaw in the tool itself. The fix is either to soften the claim to 'variable' without 'realistic,' or to report the analysis and a trajectory comparison.\n\nAlso minor: Section V-B says 'adding two new metrics' then 'new four metrics,' with a total of 32. The count works if 28 + 4 = 32, so the 'two' is likely a typo. Not a big deal, but it suggests the evaluation text wasn't carefully cross-checked.\n\nThe citation pattern looks fine. The related work covers the main alternatives, and relying on their own [1] is legitimate since the core system is there.\n\nWho is this for? Anyone building or testing human-aware navigation systems who needs scriptable, varied human behavior in simulation. It's a tool paper, not a scientific experiment, and it should be judged as that. I'd send it to review: a serious referee can ask for the missing validation details and the paper will be better for it. Verdict conditional.","headline":"Useful open-source engineering update; the realism claim outruns the evidence.","tokens_in":8691,"tokens_out":1918,"would_cite":true,"duration_ms":18408,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"HuNavSim 2.0 claims that adding controlled random variation to Social Force Model parameters—and scripting human actions with Behavior Trees—produces variable, more realistic human navigation trajectories, and bundles this with wrappers…","keywords":["social force model","behavior trees","human-aware navigation","pedestrian simulation","robot navigation benchmarking","ROS 2","navigation metrics","simulation variability"],"falsifier":"Run the same scenario twice in HuNavSim 2.0, once with default deterministic parameters and once with stochastic parameter sampling, and measure the distribution of trajectories (for example, path curvature, speed profiles, and clearance distances) against a public dataset of recorded pedestrian crossings and overtakes; if the stochastic trajectories are not statistically closer to or no more variable than the deterministic ones, the realism claim is falsified.","tokens_in":7716,"feed_emoji":"🚶","tokens_out":7949,"duration_ms":76536,"temperature":0.7,"pith_summary":"HuNavSim 2.0 tries to establish that a human-navigation simulator can produce believable, varied pedestrian behavior rather than the uniform, deterministic motion typical of crowd models. It does this by injecting controlled random variation into the parameters of the Social Force Model and by giving users a library of Behavior Tree actions and conditions to script complex human activities such as forming a conversation, inspecting a robot, or following a coworker. The paper argues that this combination, together with wrappers for several robot simulators and a 32-metric evaluation suite, makes the tool useful for developing and benchmarking human-aware robot navigation. A sympathetic reader would see the contribution as an engineering platform whose central bet is that parameter noise plus composable behaviors is enough to bridge the gap between scripted agents and real pedestrian unpredictability.","feed_headline":"Noisy crowd forces make simulated people walk more naturally","feed_subtitle":"Behavior trees and varied pedestrian motion let robot navigation be tested in realistic, diverse shared spaces.","key_machinery":"The load-bearing mechanism is the Social Force Model with parameter noise, where each agent's motion is driven by forces toward goals, away from obstacles, and between agents, and where force parameters are sampled from normal distributions ($\\mathcal{N}(\\mu,\\sigma^2)$) in feasible ranges to vary behavior across agents and runs. The second mechanism is a Behavior Tree interpreter: a tree of action nodes (GoTo, FollowAgent, ApproachRobot, ConversationFormation, SaySomething) and condition nodes (isAtPosition, IsSpeaking, IsLookingAtMe) that determines which behavior an agent executes and when to switch. The parameter sensitivity analysis is what guarantees the noise stays small enough to be realistic while large enough to be visible. The whole system is organized as a manager that receives agent states from a robot simulator wrapper, computes next states from the active behavior trees, and sends them back, with an evaluator module that logs data and computes metrics.","core_discovery":"The central discovery is an architecture-level claim: behavioral variability in simulated humans can be produced by treating the Social Force Model's parameters (such as desired speed, relaxation time, and force strengths) as random variables sampled from normal distributions over precomputed feasible ranges. With this change, repeated runs of the same scenario no longer produce identical agent trajectories; instead they produce a family of plausible trajectories. On top of this, the paper adds a set of Behavior Tree nodes that encode social conditions (e.g., is someone speaking, is another agent looking at me) and actions (approach robot, follow agent, form conversation, say something), so that whole interaction sequences can be authored as trees and executed by the HuNavSim manager. The tool's contribution is therefore twofold: a local navigation layer with tunable stochasticity and a global behavior layer with composable social scripts, both exposed through a scenario editor and measured by an extensible metric suite.","pith_inferences":["The paper does not compare noisy trajectories against real pedestrian data, so its realism claim is untested; a natural extension would be fitting the noise variances to recorded pedestrian trajectories and reporting the divergence reduction before and after.","The single sensitivity analysis that sets the feasible ranges may not transfer to other densities or environments; the tool's own scenario editor makes this easy to check by running the same behavior tree at different crowd densities and measuring trajectory variability.","Because behavior trees encode social scripts separately from the noisy motion layer, users could also author proactive social behaviors that the current example scenarios only begin to explore, such as queuing, helping, or gesturing."],"forward_implications":["If the noise injection works as claimed, every stochastic run of a scenario is a different test case, so navigation policies evaluated in HuNavSim 2.0 face a distribution of human behaviors instead of one scripted response.","Behavior Tree actions and conditions let scenario authors compose social episodes—group conversations, robot inspection, follow-me—so benchmarks can cover interaction-rich situations that plain point-to-point crowd motion cannot represent.","Because the tool wraps several robot simulators, the same behavior trees, scenario files, and metric definitions can be reused across simulators, making cross-simulator comparisons more direct.","The 32-metric evaluation suite, including the four danger-and-surprise metrics, gives users a common yardstick to compare social navigation approaches without committing to one fixed benchmark.","The interactive scenario editor turns scenario creation into a visual process, shortening the time between designing a social situation and running it in simulation."],"supporting_citations":[{"why":"Defines the original HuNavSim architecture, the metrics list, and the core manager-wrapper design that version 2.0 extends.","marker":"[1]"},{"why":"Supplies the principles and guidelines that motivate the open, extendable metric system for human-aware navigation evaluation.","marker":"[2]"},{"why":"Provides the formal Behavior Tree model that the new action and condition nodes are programmed against.","marker":"[4]"},{"why":"Introduces the Social Force Model, the underlying pedestrian motion model whose parameters are perturbed with noise.","marker":"[5]"},{"why":"Extends the Social Force Model to pedestrian social groups, which HuNavSim uses for group-aware movement.","marker":"[24]"},{"why":"Provides empirical evidence on the walking behavior of pedestrian social groups and its crowd dynamics impact, supporting the group extension.","marker":"[25]"},{"why":"Supplies the four new danger-and-surprise metrics added in HuNavSim 2.0.","marker":"[26]"}],"fun_headline_variants":["Randomized social forces make simulated crowds behave more human","Behavior trees give simulated pedestrians realistic social actions","HuNavSim 2.0 brings stochastic realism to human-robot simulation","Diverse pedestrian paths from tunable model parameters and trees"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The central claim stands or falls on the assumption that sampling the Social Force Model's parameters from bell-curve distributions—within ranges the paper does not reveal—produces trajectories closer to real human walking, since the paper provides no comparison with recorded human motion.","fun_headline_variants_meta":{"raw":{"variants":["Randomized social forces make simulated crowds behave more human","Behavior trees give simulated pedestrians realistic social actions","HuNavSim 2.0 brings stochastic realism to human-robot simulation","Diverse pedestrian paths from tunable model parameters and trees"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000231,"raw_usage":{"total_tokens":1428,"prompt_tokens":827,"completion_tokens":601,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":443,"completion_tokens_details":{"reasoning_tokens":534}},"tokens_in":443,"tokens_out":601,"duration_ms":7699,"temperature":1.0,"reasoning_tokens":534,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T14:50:55.707481+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same scenario twice in HuNavSim 2.0, once with default deterministic parameters and once with stochastic parameter sampling, and measure the distribution of trajectories (for example, path curvature, speed profiles, and clearance distances) against a public dataset of recorded pedestrian crossings and overtakes; if the stochastic trajectories are not statistically closer to or no more variable than the deterministic ones, the realism claim is falsified.","supporting_citations":[{"cited_title":"Hunavsim: A ros 2 human navigation simulator for benchmarking human-aware robot navigation,","cited_arxiv_id":null,"evidence_quote":"Defines the original HuNavSim architecture, the metrics list, and the core manager-wrapper design that version 2.0 extends."},{"cited_title":"Colledanchise and P","cited_arxiv_id":null,"evidence_quote":"Provides the formal Behavior Tree model that the new action and condition nodes are programmed against."},{"cited_title":"Social force model for pedestrian dynamics","cited_arxiv_id":null,"evidence_quote":"Introduces the Social Force Model, the underlying pedestrian motion model whose parameters are perturbed with noise."},{"cited_title":"Experimental study of the behavioural mechanisms underlying self-organization in human crowds,","cited_arxiv_id":null,"evidence_quote":"Extends the Social Force Model to pedestrian social groups, which HuNavSim uses for group-aware movement."},{"cited_title":"The walking behaviour of pedestrian social groups and its impact on crowd dynamics,","cited_arxiv_id":null,"evidence_quote":"Provides empirical evidence on the walking behavior of pedestrian social groups and its crowd dynamics impact, supporting the group extension."},{"cited_title":"Towards benchmarking human-aware social robot navigation: A new perspective and metrics,","cited_arxiv_id":null,"evidence_quote":"Supplies the four new danger-and-surprise metrics added in HuNavSim 2.0."}],"review_version":1}