{"id":"f01843ba-c9f1-425b-8782-2e0182c85d94","arxiv_id":"2509.08757","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"SocialNav-SUB introduces a VQA benchmark for social robot navigation and shows current VLMs underperform rule-based and human-agreement baselines on spatial, spatiotemporal, and social reasoning questions.","lead":"SocialNav-SUB is a new benchmark that tests whether vision-language models can understand social robot navigation scenes. It finds the best tested model still agrees with human answers less often than a simple rule-based system does.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The 0.62 vs 0.64 o4-mini/rule-based PA gap in Table 1 is not shown to be statistically robust once questions are clustered by the 60 underlying scenarios, so the claim that the best VLM underperforms the rule-based baseline is unsupported.","rationale":"SocialNav-SUB is a well-documented and useful benchmark: it releases code and data, grounds labels in a human-subject study, validates the tracking pipeline on CODa, and includes ablations. My concern is not that the benchmark is worthless but that the headline quantitative claim is over-stated. The abstract says the best VLM 'still underperforms simpler rule-based approach and human consensus baselines.' The human-consensus gap is large and likely robust, but the rule-based comparison rests on a 0.02 PA gap. Because the reported 0.01 standard errors ignore the 60-scenario structure, the evidence for that specific part of the headline is weaker than presented. I selected this concern over the reader's PHALP-tracking concern because tracking error affects the input representation for both VLM and human pipelines, and the benchmark's ground truth is defined as human agreement; in contrast, the clustering issue directly tests whether the reported numbers support the claimed ranking. The fix is straightforward: release per-scenario scores and report cluster-robust inference. This is a conditional-acceptance issue, not a rejection, so the reader's CONDITIONAL verdict should stand unchanged.","tokens_in":21451,"tokens_out":12209,"duration_ms":100715,"concrete_test":"Recompute the PA differences in Table 1 with scenario-level clustering: compute per-scene mean PA for o4-mini and for the rule-based baseline on the 60 scenarios, then run a paired permutation test (or a cluster-robust Wald test) on the 60 scene-level differences, overall and for each question category. If the 95% confidence interval includes zero, the 'underperforms rule-based' claim should be downgraded to 'no significant difference' for the best VLM.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Table 1's headline comparison is o4-mini PA 0.62 ± 0.01 versus the rule-based baseline's 0.64 ± 0.00. The reported standard errors treat the 4,968 questions as independent, but these questions are generated from only 60 SCAND scenarios, with roughly 83 questions per scenario (Appendix 7.7). Questions from the same scenario share crowd density, occlusion, environment, and question-order effects, so the effective sample size is far below 4,968. No cluster-robust standard errors, scene-level paired test, or significance test is reported. With an intra-scenario correlation of even 0.05–0.1, the design effect inflates the standard error to roughly 0.02–0.03, making the 0.02 gap non-significant. The text itself hedges by noting that the bolded VLM result 'may be statistically tied.' If the o4-mini/rule-based gap is not significant, the central claim that all tested VLMs underperform the rule-based baseline is not established; only the comparison to the human oracle (0.74 vs 0.62) would remain robust. Category-level comparisons, such as social reasoning (0.66 vs 0.71), are subject to the same clustering problem.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper introduces SocialNav-SUB, a VQA benchmark for evaluating vision-language models (VLMs) on scene understanding for social robot navigation. Built from 60 curated SCAND scenarios, it uses PHALP-based 3D tracking and Kalman smoothing to produce annotated front-view images and bird's-eye views, together with 4,968 multiple-choice questions probing spatial, spatiotemporal, and social reasoning. Human responses from a Prolific study serve as ground truth, with PA and CWPA metrics and average-human and human-oracle baselines. The authors evaluate several closed- and open-source VLMs and report that the best VLM, OpenAI o4-mini, reaches PA 0.62, below a rule-based baseline (0.64) and the human oracle (0.74). They also perform ablations on chain-of-thought and BEV prompts, a waypoint-selection experiment, and a failure-case analysis. The manuscript claims this is the first VQA benchmark for social robot navigation evaluating these reasoning dimensions against human baselines.","tokens_in":21673,"tokens_out":5363,"duration_ms":50458,"significance":"If the empirical claims hold, SocialNav-SUB would be a useful community resource: it provides a human-labeled VQA benchmark in a domain that currently lacks systematic evaluation, with multiple annotators per scenario, two agreement metrics, object-centric/BEV visual prompts, and a public-release plan. The ablation studies and the waypoint-selection validation are constructive steps toward understanding how VLMs can be integrated into social navigation stacks. The benchmark also generates falsifiable, field-relevant questions about VLM spatial and social reasoning. However, the headline comparison against the rule-based baseline is currently not statistically supported, and the rule-based baseline uses privileged tracker state rather than the visual inputs given to VLMs and humans, so the central claim needs re-analysis and reframing before the benchmark's conclusions can be accepted.","major_comments":[{"comment":"The headline comparison between o4-mini (PA 0.62 ± 0.01) and the rule-based baseline (PA 0.64 ± 0.00) treats the 4,968 questions as independent, but these questions are nested in only 60 SCAND scenarios, with about 83 questions per scenario (Appendix 7.7). Questions from the same scenario share crowd layout, occlusion patterns, environment, and question-order effects, so the effective sample size is far smaller than 4,968. No cluster-robust standard errors, scenario-level bootstrap, or paired permutation test is reported. With an intra-scenario correlation of even 0.05, the standard error on the difference would more than double, and the 0.02 gap would no longer be significant. The table caption itself concedes that the bolded VLM result 'may be statistically tied.' Because the paper's central claim is that all VLMs underperform the rule-based baseline, this comparison must be re-analyzed with scene-level clustering; the gap against the human oracle (0.74 vs 0.62) is larger and may survive, but the rule-based comparison is not established.","section":"Section 4.2, Table 1"},{"comment":"The rule-based baseline is constructed from 'the position data of pedestrians in the scene' and uses hand-crafted rules such as line-intersection checks and distance cutoffs, whereas VLMs and human annotators are given only images (front-view and BEV). This is not a like-for-like test of visual scene understanding: the rule-based baseline effectively receives privileged geometric state from the PHALP tracker. The conclusion that a 'simpler rule-based approach' outperforms VLMs is therefore misleading as stated. The comparison would be more informative if the rule-based baseline were required to operate on the same visual inputs, or if the manuscript clearly framed it as a perception-plus-rules pipeline that upper-bounds what can be done with accurate positions.","section":"Section 4.1 and Appendix 7.10"},{"comment":"The benchmark's spatial questions ask for categorical relations such as ahead, left, right, and behind, but the PHALP-based tracking pipeline reports an average displacement error of 0.67 ± 0.14 m. In crowded scenes, distances between the robot and pedestrians can be on the order of 1–2 m, so this error is large relative to the distinctions being tested. Because the same BEV positions are used to construct the VLM prompts, the human labels, and the rule-based baseline, tracking error is a shared source of label noise. The paper should quantify sensitivity to this error, for example by reporting results on a subset of scenes with high-confidence tracks below a displacement threshold, or by perturbing the BEV positions and measuring the resulting change in PA.","section":"Section 3.2 and Appendix 7.3"},{"comment":"The query protocol is described inconsistently. Section 4.1 says chain-of-thought 'provides the previous answers of the VLM for future questions,' while Section 4.2 says 'we run our experiments by querying each VLM model once per unique question.' These statements conflict: either each question is answered independently, or prior answers are fed back into the prompt. The ablation rows labeled 'No CoT' also do not clarify whether the question order and inter-question context remain the same. Because the manuscript emphasizes a fair comparison with the sequential human-subject protocol, the exact prompt construction, including whether previous model answers are included and in what order, must be specified precisely for reproducibility.","section":"Section 4.1 and Section 4.2"}],"minor_comments":[{"comment":"The text refers to 'Table 7.7' for question details, but the relevant table is Table 6 and the appendix is numbered 7.7; the cross-reference should be corrected.","section":"Section 3.3"},{"comment":"The Table 1 caption says the bolded VLM result 'may be statistically tied,' but Section 4.2 states that o4-mini 'still has a considerable gap' compared to the rule-based baseline; these statements should be made consistent with the statistical evidence after the clustering analysis is added.","section":"Table 1 and Section 4.2"},{"comment":"The rule-based baseline is described qualitatively ('cutoff values,' 'draw a line,' 'if the lines intersect'); for reproducibility, the exact thresholds, coordinate conventions, and decision rules should be provided, ideally as pseudocode or with the released implementation.","section":"Appendix 7.10"},{"comment":"The statement that PA 'is essentially the expected cosine similarity between the model's predictions and the distribution of human responses' is not literally correct unless the answer vectors are normalized; please either derive this equivalence or rephrase it as an average one-hot agreement.","section":"Section 3.4, Equation (1)"},{"comment":"The description of the prompt says the VLM receives 'the next 9 images' in addition to the image shown on the left, but the pipeline is described as a 2.5 s segment sampled at 4 Hz, which would produce 10 frames; please clarify the frame count and ordering.","section":"Appendix 7.5"}],"recommendation":"major_revision","confidential_remarks":"The benchmark itself is a valuable contribution and the concerns raised here are fixable with additional analysis rather than fatal. I recommend major revision rather than rejection because the dataset, human-subject protocol, and ablation framework are worth publishing, but the central empirical claim currently rests on a comparison that the authors themselves hedge as possibly tied statistically, and the rule-based baseline is not a visual baseline. The editors may wish to ask the authors for scene-level clustered inference and a reframed comparison with the rule-based baseline."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"SocialNav-SUB is a genuinely useful benchmark, and the release of code and data makes it a resource the field can build on. The empirical claim that the best VLM is statistically tied with or even loses to a simple rule-based baseline, however, is not as solid as the abstract suggests. The standard errors ignore clustering by the 60 underlying scenarios, so the 0.62 vs 0.64 gap is within the noise. The human-oracle gap (0.74 vs 0.62) is large and robust; the rule-based gap is not.\n\nWhat's new: this is the first VQA benchmark specifically aimed at social robot navigation scene understanding, with questions split into spatial, spatiotemporal, and social reasoning. The object-centric prompting with numbered pedestrians in both front view and BEV is a sensible design choice, and the human-subject study with 153 participants gives you a real distribution to compare against. The ablation experiments and waypoint-selection follow-up are useful extras. The authors are honest about the campus-only scenario and limited model set.\n\nSoft spots, in order:\n1. The clustering issue. With 4,968 questions drawn from 60 scenes, there's obvious within-scene correlation. The authors note in the table caption that the bolded VLM result 'may be statistically tied,' which is exactly the right worry. They need cluster-robust standard errors or a scene-level permutation test before claiming VLMs underperform the rule-based baseline. The comparison to human oracle is likely robust, but the rule-based comparison is not shown to be.\n2. The PA metric is described as 'essentially the expected cosine similarity.' For one-hot model answers, that's not right; it's simply mean agreement with human responses. Minor, but it should be fixed.\n3. The 3D tracking error of 0.67 m is nontrivial relative to the spatial categories ('ahead,' 'behind,' etc.). The paper validates on CODa but doesn't quantify how much this uncertainty propagates into the rule-based baseline or into human answers that are based on the same BEV renderings. This could matter for spatial and spatiotemporal questions.\n4. The CoT protocol gives the VLM its own previous answers as context. Humans answer the questions in sequence, but presumably without that explicit carry-forward. That asymmetry should be clarified and ideally controlled in the analysis.\n5. The scene-selection weights in Appendix 7.2 are never reported. Minor.\n\nWho this is for: anyone building or evaluating VLMs for robotics, and the social navigation community specifically. The benchmark deserves a serious referee; the statistical weaknesses are fixable and the resource itself is valuable. I'd suggest sending it to review rather than desk rejecting.","headline":"A useful new VQA benchmark for social robot navigation with reproducible code and data, but the headline VLM-vs-rule-based gap is not statistically supported once you cluster by scene; the human-oracle gap is the real finding.","tokens_in":22249,"tokens_out":3661,"would_cite":true,"duration_ms":368523,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"State-of-the-art vision-language models underperform a simple rule-based baseline and human consensus on scene understanding for social robot navigation, according to the new SocialNav-SUB benchmark.","keywords":["social robot navigation","vision-language models","visual question answering benchmark","spatial reasoning","spatiotemporal reasoning","social reasoning","bird's-eye view","human ground truth"],"falsifier":"Re-run the 60 scenarios with human answers and rule-based answers computed from a precise tracking system (for example, motion capture or multi-camera calibration) instead of the monocular tracking pipeline, and give VLMs the same corrected bird's-eye views; if their agreement rises above the rule-based baseline, the underperformance result is an artifact of the tracking errors rather than of VLM reasoning.","tokens_in":21225,"feed_emoji":"🤖","tokens_out":9287,"duration_ms":51758,"temperature":0.7,"pith_summary":"This paper introduces SocialNav-SUB, a visual question-answering benchmark that tests whether large vision-language models (VLMs) understand social robot navigation scenes. It asks spatial, spatiotemporal, and social-reasoning questions about 60 crowded real-world scenarios, with human answers serving as ground truth. The central empirical finding is that the best tested VLM (OpenAI o4-mini) agrees with humans 62% of the time, while a hand-crafted rule-based baseline reaches 64% and a human oracle reaches 74%. This matters because VLMs are being proposed as scene-understanding modules for socially compliant robot navigation, and the result suggests they are not yet reliable enough for that role. The benchmark also shows where the gap lives: spatial and spatiotemporal reasoning are the weakest areas, and social reasoning improves when spatial information is accurate.","feed_headline":"0.62 vs 0.64: vision models lag rule-based social navigation","feed_subtitle":"Best tested vision-language model agrees with humans less often than simple hand-crafted rules on crowded scenes.","key_machinery":"The carrying object is the benchmark itself: a VQA dataset plus evaluation protocol. Each item pairs a front-view RGB clip with a bird's-eye-view (BEV) image in which pedestrians are tracked by the PHALP algorithm (a monocular 3D human tracker), Kalman-smoothed, and projected from robot odometry into numbered color-coded circles; this object-centric representation is shown to both humans and VLMs. Two agreement metrics, probability of agreement (PA) and consensus-weighted probability of agreement (CWPA), score answers against the distribution of at least five human responses per question. The comparator that produces the headline result is a simple rule-based baseline that answers from the same pedestrian position data using hand-crafted cutoffs and line-intersection tests. Chain-of-thought prompting is the query mechanism that links questions sequentially, and ablations of BEV and CoT show their contributions.","core_discovery":"SocialNav-SUB is the first VQA benchmark built specifically for social robot navigation scene understanding. It provides 4,968 human-labeled multiple-choice questions derived from 60 SCAND scenarios, each represented as a 2.5-second multi-view clip with numbered, color-coded pedestrians in front-view and bird's-eye-view images. Across the tested VLMs—Gemini 2.0 and 2.5, GPT-4o, OpenAI o4-mini, and LLaVa-Next-Video—the best model, o4-mini, achieves a probability of agreement (PA) of 0.62, below the rule-based baseline's 0.64 and the human oracle's 0.74, while average human PA is 0.60. The largest shortfalls are in spatial and spatiotemporal reasoning, while social reasoning comes closest to human-level agreement. The paper also reports that chain-of-thought prompting improves social reasoning and that supplying ground-truth spatial answers improves social-reasoning performance, indicating that poor spatial grounding is a bottleneck.","pith_inferences":["If the paper's monocular tracking has systematic rather than random error, the human labels and rule-based baseline may be biased in the same direction; re-scoring with ground-truth tracks could narrow or eliminate the VLM gap. This is a testable reconstruction, not a claim the paper makes.","Because human annotators often disagree on the same scene, exact-match agreement may understate a VLM's competence on genuinely subjective questions; a distributional scoring scheme could tell whether remaining errors are wrong or merely one plausible reading.","The same BEV-plus-numbered-pedestrian VQA template could be ported to other dynamic robot domains, such as driving or drone navigation, to test whether the observed underperformance is specific to social scenes or a general deficit in spatiotemporal reasoning.","Making the human labels publicly available turns the benchmark into training material, not just an evaluation: fine-tuning or reinforcement learning from these answers could plausibly close part of the gap that zero-shot prompting has not closed."],"forward_implications":["A robot using current VLMs as its only scene-understanding module would agree with human judgments less often than a robot using the simple rule-based baseline, so near-term deployment should pair VLMs with dedicated perception and reasoning modules.","Because supplying correct spatial and spatiotemporal answers to the VLM raised its social-reasoning performance, spatial grounding is a bottleneck: improving spatial and dynamic perception should improve higher-level social inference.","Query format materially changes measured capability: chain-of-thought reliably helps social reasoning, while BEV overlays help some models and not others, so reported VLM performance is tied to the prompt interface.","The waypoint-selection experiment links scene understanding to action: when social scene context came from the human oracle rather than from random or model-generated context, all evaluated VLMs selected the human operator's waypoint more often.","Failure rates are uneven: models err more often at blind corners and indoors and on 'yielding to' and 'overtaking' labels, identifying targeted environments and action categories for future data collection or fine-tuning."],"supporting_citations":[{"why":"Supplies the 60 crowded real-world social navigation scenarios from which all benchmark questions are drawn.","marker":"[4]"},{"why":"Produces the monocular 3D pedestrian tracks that are projected into the bird's-eye views, determining the spatial information both humans and VLMs see.","marker":"[46]"},{"why":"Is the earlier VLM-based social navigation work whose small, controlled evaluations SocialNav-SUB extends.","marker":"[9]"},{"why":"Provides prior evidence that VLMs lack robust spatial cognition, the gap the benchmark is designed to quantify.","marker":"[10]"},{"why":"Is one of the closed-source general-purpose VLMs benchmarked, contributing to the main performance table.","marker":"[6]"},{"why":"Is the reasoning-capable Gemini family evaluated; its scores populate the comparison table.","marker":"[7]"},{"why":"Is the best-performing VLM in the paper; its 0.62 probability of agreement is the headline number compared against the baselines.","marker":"[16]"},{"why":"Is the open-source video VLM included as the deployable, locally runnable baseline.","marker":"[17]"},{"why":"Supplies the human-agreement evaluation approach that the PA and CWPA metrics adapt.","marker":"[51]"}],"fun_headline_variants":["VLMs underperform rule-based methods in social navigation benchmark","SocialNav-SUB: best VLM hits 0.62, below rule-based 0.64","Spatial and spatiotemporal reasoning drag down VLMs on social nav","New VQA benchmark exposes VLM gaps in social robot navigation"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The ground-truth human answers and the rule-based baseline are both built from estimated pedestrian positions in the bird's-eye view, produced by monocular tracking with an average displacement error of 0.67 meters, so if those positions are systematically wrong, the benchmark's comparison is distorted.","fun_headline_variants_meta":{"raw":{"variants":["VLMs underperform rule-based methods in social navigation benchmark","SocialNav-SUB: best VLM hits 0.62, below rule-based 0.64","Spatial and spatiotemporal reasoning drag down VLMs on social nav","New VQA benchmark exposes VLM gaps in social robot navigation"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00025,"raw_usage":{"total_tokens":1613,"prompt_tokens":1062,"completion_tokens":551,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":678,"completion_tokens_details":{"reasoning_tokens":470}},"tokens_in":678,"tokens_out":551,"duration_ms":5408,"temperature":1.0,"reasoning_tokens":470,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T15:59:14.592271+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the 60 scenarios with human answers and rule-based answers computed from a precise tracking system (for example, motion capture or multi-camera calibration) instead of the monocular tracking pipeline, and give VLMs the same corrected bird's-eye views; if their agreement rises above the rule-based baseline, the underperformance result is an artifact of the tracking errors rather than of VLM reasoning.","supporting_citations":[{"cited_title":"Rajasegaran, G","cited_arxiv_id":null,"evidence_quote":"Produces the monocular 3D pedestrian tracks that are projected into the bird's-eye views, determining the spatial information both humans and VLMs see."},{"cited_title":"Openai o3 and o4-mini system card, 2025","cited_arxiv_id":null,"evidence_quote":"Is the best-performing VLM in the paper; its 0.62 probability of agreement is the headline number compared against the baselines."},{"cited_title":"Zhang, B","cited_arxiv_id":null,"evidence_quote":"Is the open-source video VLM included as the deployable, locally runnable baseline."}],"review_version":2}