{"id":"1dc71080-cf95-45aa-a382-71e8595f6ed2","arxiv_id":"2607.09261","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Matched real-world vs VR co-navigation with a PR2 shows similar social-awareness ratings and consistent trajectory and head-orientation patterns, supporting VR for multimodal SRN studies.","lead":"A within-subjects study (N=21) found that people rate a PR2 robot’s social awareness similarly in VR and the real world, and that trajectories plus head orientations show consistent patterns across matched co-navigation scenarios. This supports using VR as a safer, cheaper platform for multimodal social-navigation experiments before real-robot trials.","discovery_kind":"extension","skeptic_critique":{"model":"grok-4.5","headline":"Descriptive pattern-matching without equivalence tests or selection criteria cannot underwrite the claim that VR “preserves” multimodal dynamics.","rationale":"The reader correctly isolates the weakest link: descriptive similarities plus a shared planner are treated as sufficient for a reliability claim even though objective metrics differ systematically, questionnaires block direct comparison, and the authors themselves flag the analysis as preliminary. No stronger internal inconsistency or fabrication is present; the design is clean and the honesty about limitations is a credit. The concern therefore does not overturn the CONDITIONAL verdict but confirms it: stronger equivalence statistics and transparent trajectory selection remain the necessary conditions for the central claim to hold. My concrete test operationalizes exactly that missing step.","tokens_in":13115,"tokens_out":569,"duration_ms":5729,"concrete_test":"Release the full trajectory set (or the exact selection rule for the 240 interactions) and recompute Tables 2–4 with paired equivalence tests (e.g., TOST on velocity, min-distance, path deviation, and head-cosine WA/AC) using a pre-declared ±10–15% margin of the RW means. If any primary metric fails equivalence or the selected subset differs materially from the full set, the “preserves multimodal dynamics” claim weakens.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The central claim (Abstract, §5.4, §7) that VR “preserves the multimodal interaction dynamics” and is therefore a “reliable platform” rests on (i) identical medians for Q1/Q2 despite r=0.17, (ii) qualitative similarity of velocity/jerk/min-distance/path-deviation/head-cosine tables (Tables 2–4), and (iii) the same CoHAN planner + PR2 morphology. Yet the paper itself reports systematic differences (humans slower, larger min-distance, lower jerk in VR; naturalness median=3) and states that non-identical questionnaires preclude statistical comparison (§5) while calling the analysis “preliminary” and needing “deeper statistical validation” (§6). No equivalence bounds, confidence intervals, or pre-specified similarity criteria appear; the 240 trajectories are a post-hoc subset of the 840 possible rounds with no selection rule given. Consequently the leap from “descriptively similar trends” to “preserves … reliable” is under-supported by the evidence the authors present.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.5","summary":"The paper develops a Unity/ROS VR prototype that replicates a motion-capture arena and a PR2 robot controlled by the CoHAN socially-aware planner (plus a simple gaze behavior). It reports a counterbalanced within-subjects study (N=21) comparing two co-navigation scenarios (orthogonal crossing, pass-by) in real-world (RW) versus VR. Subjective Likert items assess perceived social awareness and comfort; objective metrics include velocity/jerk at closest approach, minimum human-robot distance, path deviation from shortest path, and head-orientation cosine distances (toward future path and toward the robot). The authors conclude that participants perceive the robot similarly and that VR captures locomotion and head-orientation patterns consistent with RW, so VR is a reliable platform for multimodal socially-aware navigation studies.","tokens_in":13377,"tokens_out":1123,"duration_ms":19170,"significance":"A carefully matched RW-VR comparison that includes head orientation (not only planar trajectories) would be a useful methodological contribution for the SRN/HRI community, where real-world multimodal data collection is costly and hard to control. The shared robot morphology, identical planner, and dual-scenario design are strengths; the open acknowledgment of limitations (audio, FOV, naturalness) is also welcome. If the similarity claim holds under tighter statistical scrutiny, the platform could support safer preliminary studies and richer multimodal data collection before real-robot deployment.","major_comments":[{"comment":"Abstract, §5.1 and §7: The central claim that participants “perceive the robot’s socially aware navigation similarly” and that VR “preserves the multimodal interaction dynamics” rests on identical medians for Q1/Q2 (both 4) despite a reported Pearson r = 0.17 and greater variance in VR. Non-identical questionnaires are explicitly noted as precluding deeper statistical comparison (§5). Without equivalence tests, confidence intervals, or pre-specified similarity criteria, the leap from “same median” to “similarly / preserves / reliable” is under-supported.","section":"§5.1, Abstract, §7"},{"comment":"Tables 2–4 and §5.4: Systematic differences appear alongside the claimed consistencies—humans move slower, exhibit lower jerk, and maintain larger minimum distances in VR; path-deviation and head-cosine patterns are directionally similar but not statistically tested for equivalence. The 240 trajectories are described as a post-hoc subset of the collected rounds with no selection rule or inclusion criteria given. These differences and the missing selection protocol weaken the assertion that VR “captures human interaction behaviors in ways consistent with real-world observations.”","section":"Tables 2–4, §5.4"},{"comment":"Table 4 and §5.4: Cosine-distance calculations use different FOV thresholds (90° in VR vs. assumed 150° in RW) and set out-of-FOV values to zero. This ad-hoc asymmetry can itself produce the reported numerical similarity; a sensitivity analysis or identical FOV treatment is needed before claiming that head-orientation patterns are preserved.","section":"Table 4, §5.4"},{"comment":"§6: The authors themselves label the analysis “preliminary” and call for “deeper statistical validation” with larger samples. Given that the paper’s title and abstract already assert validation and reliability, either the statistical treatment must be strengthened (equivalence bounds, mixed-effects models, trajectory-selection protocol) or the claims must be substantially tempered to match the evidence actually presented.","section":"§6"}],"minor_comments":[{"comment":"§4.3 / Table 1: Q1 is administered only after RW and Q2–Q4 only after VR; the direct-comparison items (Q5–Q6) are post-hoc. A fully parallel instrument would have allowed paired tests and should be noted as a design limitation.","section":"§4.3"},{"comment":"Fig. 5: Representative trajectories are helpful, but error bands or density plots across the 60 interactions per cell would better convey variability.","section":"Fig. 5"},{"comment":"§3: The pitch-to-velocity mapping (max 1.5 m/s) and avatar animation details are free parameters that affect naturalness ratings (median 3); a brief sensitivity note would help readers assess generalizability.","section":"§3"},{"comment":"Minor typographical inconsistencies appear (e.g., “Weconductedacomparativeuserstudy”, spacing around citations). A careful proof-read is needed.","section":null}],"recommendation":"major_revision","confidential_remarks":"The work is a competent empirical methods paper rather than a algorithmic advance; it is appropriate for an HRI or robotics methods venue but may be borderline for a top general robotics journal unless the statistical support is materially strengthened. The low r=0.17 and systematic speed/distance offsets are the load-bearing weaknesses; if the authors can add equivalence tests or clearly qualify the claims, the contribution becomes publishable."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The punchline is simple: this is a clean within-subjects N=21 comparison of the same PR2 + CoHAN setup in a motion-capture arena and its Unity replica, recording both trajectories and head orientation for orthogonal crossing and pass-by. That matched multimodal dataset is the actual new piece; prior VR-SRN work already looked at proxemics, velocity, and faces, but not this head-to-head with head cues.\n\nThey do several things well. Design is appropriate (counterbalanced order, habituation, fixed planner and morphology so the robot is a constant stimulus). Tables 2–4 and Fig. 5 show real descriptive consistencies in path deviation, head-cosine alignment with future path, and robot velocity profiles. The limitations section is honest about missing audio, FOV, naturalness (median 3), and the need for deeper stats. Citations cover the right prior art; no circularity or invented entities. CoHAN and the gaze planner are just fixed stimuli.\n\nThe soft spot is proportionate but real. Social-awareness ratings share a median of 4 yet correlate only r=0.17. Humans are systematically slower, keep larger min-distance, and show lower jerk in VR. Questionnaires are non-identical, so no paired stats. The 240 trajectories are a post-hoc subset of ~840 rounds with no selection rule stated. They themselves call the analysis “preliminary.” Descriptive pattern-matching plus identical medians does not underwrite the abstract/conclusion claim that VR “preserves the multimodal interaction dynamics” and is therefore “reliable.” Free parameters (controller mapping, FOV thresholds, gaze trigger) are minor but present.\n\nThis is for people already doing SRN user studies who want a practical VR platform and early multimodal signals. It is not a theoretical advance. The data and design are solid enough that a serious editor should send it to referees; they will demand equivalence tests, selection criteria, and toned-down claims. I would read the camera-ready and cite the comparative numbers if I am building a similar VR setup. Worth engaging, not a desk reject.","headline":"Useful matched RW–VR multimodal dataset for SRN, but the leap from descriptive trends to “preserves dynamics / reliable platform” is under-supported.","tokens_in":13981,"tokens_out":524,"would_cite":true,"duration_ms":9888,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Virtual reality reproduces human locomotion and head-orientation patterns with a socially aware robot closely enough to serve as a reliable study platform.","keywords":["Virtual Reality","Socially Aware Robot Navigation","Multimodal HRI","Head Orientation","Co-navigation","Human-Robot Interaction","Trajectory Analysis"],"falsifier":"A larger within-subjects study using identical questionnaire items and formal equivalence or mixed-model tests that finds significant differences in social-awareness ratings or systematic mismatches in head-orientation correlation with human–robot distance between VR and the real world would falsify the claim.","tokens_in":13997,"feed_emoji":"🥽","tokens_out":819,"duration_ms":18034,"temperature":0.7,"pith_summary":"This paper asks whether immersive virtual reality can stand in for real-world experiments when people share space with a robot that navigates in a socially aware way. Earlier VR work on social navigation mostly tracked planar paths and distances; it was unclear whether richer cues such as head orientation would match real behavior. In a within-subjects study of 21 participants walking with the same PR2 robot and planner in a motion-capture arena and its VR replica, people rated the robot’s social awareness similarly, and their trajectories and head directions followed consistent patterns in both orthogonal-crossing and pass-by encounters. The result matters because it would let researchers run safer, more controllable, and richer multimodal studies without building a full physical arena for every condition, speeding work on robots that coordinate using gaze and motion rather than pure geometry.","feed_headline":"VR matches real-world human-robot co-navigation","feed_subtitle":"Trajectories and head orientations stay consistent across VR and a physical arena in a 21-person study","key_machinery":"A matched within-subjects VR prototype that replicates a motion-capture arena and a PR2 driven by the same CoHAN socially aware planner and gaze behavior; comparison of social-awareness ratings with quantitative trajectory metrics (velocity, jerk, path deviation, minimum distance) and head-orientation cosine distances while approaching and after crossing.","core_discovery":"Participants perceive a PR2 robot’s socially aware navigation similarly in immersive VR and in the real world, and VR captures human locomotion trajectories and head-orientation cues in ways consistent with real-world co-navigation for orthogonal-crossing and pass-by scenarios, supporting VR as a reliable platform for multimodal socially aware navigation research.","pith_inferences":["If depth-perception and speed biases in VR are systematically calibrated, trajectory datasets from VR could be mixed with real-world logs for training human-motion predictors.","Validating head orientation as a proxy for attention opens tests of whether robots that condition plans on estimated human gaze improve comfort more than trajectory-only planners.","The same matched VR–real protocol could benchmark whether different robot morphologies or planners preserve cross-setting consistency of human head behavior."],"forward_implications":["Researchers can collect richer multimodal navigation data (trajectories plus head orientation) under controlled conditions without always needing a physical arena.","Preliminary user studies of multimodal socially aware navigation strategies can be run in VR before full real-world deployment.","The same robot planner and embodiment can be tested for human responses across VR and real settings with comparable perceived social awareness.","Future framework extensions (spatial audio, full-body pose, eye tracking) can build on the validated locomotion and head-orientation baseline."],"fun_headline_variants":["VR matches real-world robot co-navigation behaviors","Participants rate social robot awareness same in VR and arena","Trajectories and head cues stay consistent in VR co-navigation","VR proves reliable platform for multimodal social HRI tests","Orthogonal and pass-by robot interactions hold up in VR"],"cache_read_input_tokens":128,"weakest_assumption_plain":"That similar median ratings and consistent descriptive patterns in speed, path deviation, and head direction are enough to treat VR as preserving multimodal dynamics, even though absolute human speeds, distances, jerk, and naturalness differ and the questionnaires were not identical.","fun_headline_variants_meta":{"raw":{"variants":["VR matches real-world robot co-navigation behaviors","Participants rate social robot awareness same in VR and arena","Trajectories and head cues stay consistent in VR co-navigation","VR proves reliable platform for multimodal social HRI tests","Orthogonal and pass-by robot interactions hold up in VR"]},"model":"grok-4.5","effort":"low","cost_usd":0.003946,"raw_usage":{"total_tokens":1235,"prompt_tokens":770,"num_sources_used":0,"completion_tokens":62,"cost_in_usd_ticks":39460000,"prompt_tokens_details":{"text_tokens":770,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":403,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":770,"tokens_out":62,"duration_ms":5399,"temperature":1.0,"reasoning_tokens":403,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T04:18:59.115479+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"A larger within-subjects study using identical questionnaire items and formal equivalence or mixed-model tests that finds significant differences in social-awareness ratings or systematic mismatches in head-orientation correlation with human–robot distance between VR and the real world would falsify the claim.","supporting_citations":[],"review_version":1}