{"id":"63426dca-634d-4739-9e7b-8cce8b82fd70","arxiv_id":"2608.08077","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Embodied VLMs in an emergency evacuation benchmark often choose gathering points from pretrained textual priors rather than from rooms they visited, and their spatial memory does not follow human forgetting patterns.","lead":"This paper benchmarks seven vision-language models in a simulated emergency evacuation, asking whether their chosen gathering points are grounded in explored evidence or in pretrained text biases. It finds that model decisions frequently follow textual priors, that low light degrades spatial mapping, and that model memory does not match human memory curves.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The flagship inventory-free intervention baseline is reported as 88.6% (Table 4) but 73.8% (Table 2) and 59.5% (Table 3) for the same condition, so the headline effect size for RQ1 is not internally consistent.","rationale":"The paper's intervention design is genuinely strong: equalising textual descriptors (B1-B4), anonymising room names, and ablating the inventory are the right controlled manipulations to test whether focal-point choices are driven by textual priors. The evidence that Qwen's Conference preference is sharply reduced without the inventory is, if accurate, the paper's most convincing support for RQ1. The concern raised here is not about the validity of the intervention concept but about the numerical consistency of the headline statistic. Table 4's 88.6% baseline contradicts the same-condition baselines in Tables 2 and 3. The reader's weakest_assumption focused on JSON-as-memory (RQ3), which is a different axis; however, the reader's rationale did flag incomplete-run reporting in several tables, which partially covers this issue. Our specific check sharpens that flag: the inconsistency is in the flagship RQ1 table, not just the memory metrics. We do not think this warrants rejection, because the qualitative conclusion is likely robust and the paper is otherwise transparent; but the paper should not be accepted without reconciling these counts, because the abstract's central quantitative claim is presently unverifiable from the paper's own tables. The CONDITIONAL verdict remains appropriate, and our recommendation is 'unchanged' relative to the reader's judgment, with the explicit condition that the baseline reconciliation be completed.","tokens_in":27616,"tokens_out":9993,"duration_ms":98540,"concrete_test":"From the public repository (github.com/gabrielelamalfakcl/emrd), extract the raw terminal Term(meeting_point=...) choices for Qwen3-VL-32B Baseline runs and recompute the aggregate focal-point distribution by starting room. Compare the resulting counts against Table 2, Table 3, and Table 4. Then rerun the Inventory-Free intervention using the same harness and recompute the Conference-vs-other 2x2 table and Pearson chi-square with the reconciled baseline. If Table 4's 88.6% cannot be reproduced from the logs, the paper must correct Tables 2-4 and the abstract's '88.6% to 27.5%' claim, or explain which runs were excluded.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central causal evidence for RQ1 is the Qwen3-VL-32B inventory-free intervention: removing the object inventory drops Conference Room preference from 88.6% to 27.5% (chi-square 62.4, p<1e-13; Table 4 and cited in the abstract). This baseline is not reproducible from the paper's own tables. Table 4 ('With Inventory', n=79) reports Conference 88.6%, Kitchen 3.8%, Entrance 7.6%, Office 0%. Table 2's Baseline row for Qwen (n=79) reports Conference 73.8%, Entrance 25.0%, Other 1.2%. Table 3's Baseline row (n=79) reports Conference 59.5%, Entrance 5.1%, Other 34.2%. These are three mutually incompatible distributions for the same model and condition. If the true baseline is 73.8% or 59.5% rather than 88.6%, the qualitative conclusion (textual inventory shapes focal-point choice) may survive, but the headline magnitude, the chi-square value, and the strength of the 'post-hoc rationalisation' claim are all quantitatively wrong as stated. The paper gives no explanation for the discrepancy; it cannot be attributed to different n because all three rows report n=79. This is a load-bearing measurement inconsistency, not a stylistic issue.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper extends the Theory of Space (ToS) framework into a safety-critical, goal-driven pipeline called Explore, Map, Remember, Decide (EMRD), and uses it to evaluate seven vision-language models (VLMs) in a partially observable 3D office evacuation task. The authors introduce metrics for exploration coverage and efficiency, spatial fidelity (positional accuracy and temporal belief stability), memory persistence (Ebbinghaus decay, serial position effect, Miller's law), and cognitive decision-making (focal-point spatial grounding, FPSG). Three research questions are addressed: whether focal-point decisions are grounded in spatial evidence or driven by pre-trained semantic biases (RQ1), whether low-light or texture/color perturbations degrade spatial reasoning (RQ2), and whether VLM memory aligns with human cognitive laws (RQ3). The main claims are that VLMs frequently choose evacuation points from textual priors without spatial grounding, that low light but not texture/color tampering degrades performance, and that VLM memory does not follow human-like forgetting patterns.","tokens_in":27963,"tokens_out":3738,"duration_ms":37999,"significance":"If the central claims hold, the paper makes a useful contribution by providing a structured benchmark for evaluating embodied VLMs in safety-critical scenarios, and by highlighting that high inter-agent consensus on a focal point can coexist with weak physical grounding. The controlled prompt interventions in RQ1 are a genuine strength: removing the object inventory and equalizing room descriptors provides an external, non-circular probe of the textual-prior hypothesis, and the reported chi-square tests give quantitative support for the claim that inventory drives focal-point preference. The FPSG metric, which separates 'visited the chosen room' from 'retained the chosen room's objects,' is a practical addition to the evaluation toolbox. However, the present version contains a load-bearing internal inconsistency in the baseline measurements for the flagship intervention, and the RQ3 memory conclusions rest on an unvalidated assumption that the JSON cognitive map faithfully externalizes internal memory. These issues must be resolved before the quantitative claims can be trusted.","major_comments":[{"comment":"Table 4 reports the Qwen3-VL-32B baseline Conference Room preference as 88.6% (n=79), Table 2 reports 73.8% for the same model and baseline condition (n=79), and Table 3's Baseline row reports 59.5% (n=79). These three distributions for an identical condition are mutually incompatible, and the abstract and RQ1 rely on the 88.6% to 27.5% drop (chi-square=62.4, p<1e-13). Please reconcile these numbers or correct the claims; if the true baseline is 59.5% or 73.8%, the headline effect size, the chi-square value, and the strength of the 'post-hoc rationalisation' conclusion are not as stated.","section":"Results, Tables 2-4"},{"comment":"All Remember-phase metrics -- Time-to-Forget, Ebbinghaus memory strength, Miller's-law capacity, and FPSG -- treat objects present in the step-by-step JSON Cognitive Map as the agent's internal memory trace. No evidence is provided that this JSON output is a faithful externalisation of the model's spatial memory rather than a format-compliance artefact, a summarisation of the prompt inventory, or a sliding-window restatement of the conversation context. Without validation of this measurement assumption, or a matched human baseline, the RQ3 conclusion that 'VLM memory fundamentally diverges from human cognition' is not directly supported.","section":"Methodology, 'Prompting and cognitive mapping'; Appendix A"},{"comment":"The Ebbinghaus memory strength parameter EFC_X is estimated via MLE from the same Time-to-Forget data that is then evaluated for goodness of fit (R^2 against a Kaplan-Meier curve). A high R^2 therefore only confirms that the exponential form is a reasonable description of the pooled forgetting events; it does not by itself establish human-like biological decay. Similarly, the negative R^2 values for GPT-5.4 and Pixtral-12B are described as evidence of 'rigid, non-biological filters,' but the same pattern would arise from any non-exponential retention process, including output-formatting inconsistencies. Please test the exponential assumption against alternative distributions, or reframe the claims as 'retention is approximately exponential for some models' without the human-cognition interpretation.","section":"Remember: Memory Persistence, Eqs. (7)-(8); Table 6"}],"minor_comments":[{"comment":"Table 5 and Table 11 report conflicting values for Claude Opus 4.8 and InternVL3-14B: for example, Table 5 lists Claude Opus 4.8 Baseline AEC=28.5±10.5 and TE=23.5±13.1, while Table 11 swaps these as AEC=23.5±13.1 and TE=28.5±10.5; please align the two tables.","section":"Tables 5 and 11"},{"comment":"Table 3 and Appendix Table 9 appear to duplicate the same controlled-intervention results with identical p-values; consider presenting the full table once and referencing it from the main text.","section":"Table 3 and Appendix Table 9"},{"comment":"The Neutral condition in Table 3 has n=19 while all other rows have n=79-80; the substantially lower sample size is not explained in the text.","section":"Table 3, Neutral row"},{"comment":"The conversation context is explicitly truncated to the last 10 dialogue turns, which is itself a strong memory confound for the Miller's-law analysis; state explicitly how this context-window truncation interacts with the measured forgetting thresholds.","section":"Appendix C, 'Details on the VLMs' Hyperparameters'"}],"recommendation":"major_revision","confidential_remarks":"The inconsistency in the Qwen3-VL-32B baseline across Tables 2, 3, and 4 is severe enough that I would want to see the corrected data and a re-executed statistical analysis before recommending acceptance. I also recommend that the authors report the raw per-run focal-point and JSON-map data, or otherwise make the reproducibility artifacts explicit, since the memory claims depend on properties of the JSON outputs that cannot be checked from the current tables."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a serious look. The genuinely new thing here is the EMRD pipeline and the FPSG metric, plus the application of Ebbinghaus, serial-position, and Miller-law metrics to VLM spatial maps in a goal-driven evacuation task. The controlled intervention on Qwen3-VL-32B — removing the object inventory from the prompt — is the strongest part. It gives external evidence for the RQ1 claim that focal-point choices follow textual priors, not explored evidence. That part is not circular: they probe the mechanism with ablations, and the code and data are on GitHub.\n\nThe paper also does some things carefully. The POMDP formalisation is honest, the coverage-anchored metrics avoid survivorship bias, and the exploration/map metrics are sensible. The RQ2 finding that low light hurts but texture and colour tampering do not is a clean, falsifiable result, even if the environment is small.\n\nThe soft spots. The stress-test note is correct and load-bearing. Table 4 reports the inventory-free baseline as 88.6% Conference for Qwen; Table 2 reports 73.8% and Table 3 reports 59.5% for the same condition, all n=79. Those are mutually incompatible distributions. The 88.6-to-27.5 drop and the chi-square of 62.4 are the paper's headline numbers. The qualitative conclusion may survive, but the magnitudes as stated are unreliable. That needs an explanation and a rerun.\n\nSecond, the memory metrics treat the step-by-step JSON cognitive map as a faithful externalisation of internal spatial memory. Time-to-Forget, Ebbinghaus strength, Miller capacity, and FPSG all come from object presence and coordinates in that JSON. The paper never validates that output against a matched human baseline or any independent measure of what the model retains. So the RQ3 claim that VLM memory fundamentally diverges from human cognition is really about the JSON output under a specific prompt, not necessarily about internal memory. That is a load-bearing assumption, not a minor caveat.\n\nMinor issues: several tables have incomplete runs (Pixtral and InternVL missing data), and the single four-room office with static hazards is a narrow base for 'fundamental limitations'. Those are fixable.\n\nWho this is for: embodied-AI and VLM-safety researchers who want a concrete benchmark with metrics they can run. It deserves a serious referee — the core idea is timely and the controlled intervention is valuable — but it needs a major revision: reconcile the baseline numbers, validate or reinterpret the memory metrics, and soften the generality claims.\n\nRecommendation: send to peer review; require the revision before acceptance.","headline":"A useful safety-critical VLM benchmark with a genuine controlled intervention for RQ1, but the headline effect size is internally inconsistent and the memory metrics rest on an unvalidated JSON-as-memory assumption.","tokens_in":28461,"tokens_out":3051,"would_cite":true,"duration_ms":29457,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that in safety-critical evacuation tasks, vision-language models decide where to gather based on pre-trained semantic priors rather than on rooms they explored and mapped, and their memory does not follow human…","keywords":["vision-language models","embodied AI","spatial memory","safety-critical decision-making","focal-point","partial observability","cognitive map","spatial grounding"],"falsifier":"Swap the semantic labels in the same four-room layout, relabeling the physical Conference Room as 'Storage Closet' and the Entrance as 'Conference Room' while keeping the inventory descriptions fixed to the visual contents. If models still choose the physically larger central room, the prior is visual or geometric; if they switch to the Entrance, the decision tracks the text labels, confirming the paper's claim. A complementary check is to compare the JSON map with a free-recall probe after the episode, without asking for JSON.","tokens_in":27438,"feed_emoji":"🚨","tokens_out":7561,"duration_ms":71409,"temperature":0.7,"pith_summary":"The paper asks whether vision-language models can be trusted for safety-critical navigation such as emergency evacuation. It introduces a four-phase evaluation, Explore, Map, Remember, Decide, run on seven models in a partially observable 3D office. Its central finding is that focal-point choices, where to gather, are frequently made before or without visiting the chosen room, and controlled prompt changes show the rationale is post-hoc rationalisation: removing the room inventory drops the conference-room preference from 88.6% to 27.5%. It also finds that low light degrades spatial mapping while texture and colour changes do not, and that none of the models reproduce the human U-shaped primacy-recency retention curve. The matter because if true, current vision-language models cannot be assumed to base life-safety decisions on physical evidence.","feed_headline":"Vision-language models pick evacuation rooms from text bias","feed_subtitle":"A four-part benchmark shows models pick the same rooms from any start and that low light, not color changes, breaks mapping.","key_machinery":"The load-bearing instrument is the step-by-step Cognitive Map: the model is prompted at every step to output a JSON object with predicted 2D coordinates of every object it believes it has seen, after a deduplication check, and the paper treats this JSON as the model's externalised spatial belief state. Around this state it builds the EMRD pipeline, with exploration coverage and temporal efficiency, positional accuracy and temporal belief stability, psychological memory metrics, and the new Focal-Point Spatial Grounding score, which measures how many objects from the chosen evacuation room are actually present in the final map, anchored against total explored objects. The controlled prompt interventions, neutralising room descriptions, renaming rooms A-D, and removing the inventory, are the mechanism that isolates semantic priors from visual evidence.","core_discovery":"The central claim is negative: under partial observability, the evacuation decisions of current vision-language models are largely determined by pre-trained textual priors rather than by the spatial evidence gathered during exploration. The paper supports this with Focal-Point Spatial Grounding metrics showing that several models select a room they never entered, up to 47.9% of episodes for one model, and with an inventory-free controlled intervention in which removing the object inventory from the system prompt collapses a dominant 88.6% conference-room preference to 27.5%, with a chi-square p-value below 1e-13. The same intervention shows a recency and proximity bias: without the textual inventory, 50.6% of choices default to the room where the agent spawned. A second set of findings concerns memory: no model produces the human U-shaped serial-position curve, and forgetting patterns are either exponential with model-specific time constants or, for two models, not exponential at all. A third finding is asymmetric robustness: mapping degrades sharply under reduced visibility but is largely unaffected by texture and colour randomisation.","pith_inferences":["Editorial extension: the inventory-free result suggests a testable generalisation, that any semantic label in the prompt, not only room names, can anchor focal points; one could vary object lists across otherwise identical scenes to map each model's prior dictionary.","Editorial extension: the finding that texture and colour tampering barely hurts mapping hints that these models rely on geometric layout and object contours rather than surface appearance; a direct test would use silhouette-preserving monochrome scenes.","Editorial extension: because the JSON map is the only memory trace, the EMRD pipeline could be inverted into a training signal, rewarding models for high spatial grounding and stable object retention, rather than serving only as a benchmark."],"forward_implications":["Evacuation protocols built on current vision-language models should not treat a model's stated rationale as evidence-grounded; the same room may be selected from any starting point without visitation.","The textual inventory in the prompt is a decision lever: removing it fragments choices and shifts them to visually salient rooms or the spawn room, so prompt design, not just perception, controls the outcome.","Low-light emergencies are the dangerous failure mode: coverage, positional accuracy, and map stability all degrade under reduced visibility, while texture and colour alterations are harmless; systems will need auxiliary sensing in smoke or power failure.","Human-agent memory alignment cannot be assumed: two models show non-exponential forgetting, and none show the human U-shaped primacy-recency curve; interaction design should not rely on human-like recall.","Accurate mapping does not guarantee reliable decisions: even models with higher coverage can choose unvisited rooms."],"supporting_citations":[{"why":"Supplies the belief-probing protocol that this paper extends into the EMRD pipeline.","marker":"Zhang et al. 2026"},{"why":"Defines focal-point decisions, the decision target under study.","marker":"Schelling 1981"},{"why":"Provides the ProcTHOR 3D office environment used for navigation and evacuation experiments.","marker":"Deitke et al. 2022"},{"why":"Gives the Procrustes alignment used to compute positional accuracy of the cognitive map.","marker":"Schönemann 1966"},{"why":"Defines the forgetting curve used to test whether memory decay matches biological patterns.","marker":"Ebbinghaus 1913"},{"why":"Defines the serial-position effect used to test primacy and recency in retention.","marker":"Murdock Jr 1962"},{"why":"Defines the working-memory capacity law used for the Miller metric.","marker":"Miller 1956"}],"fun_headline_variants":["VLMs pick evacuation rooms from text bias, not maps","VLMs ignore spatial evidence in evacuation decisions","Text bias, not maps, guides VLM escape choices","Low light breaks VLM spatial maps; color does not","VLM memory diverges from human patterns, study shows"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that the step-by-step JSON cognitive map is a faithful externalisation of what the model actually remembers, rather than a format-compliant output that happens to satisfy the prompt; if the JSON is only prompt-formatting behaviour, the memory and grounding metrics do not measure memory.","fun_headline_variants_meta":{"raw":{"variants":["VLMs pick evacuation rooms from text bias, not maps","VLMs ignore spatial evidence in evacuation decisions","Text bias, not maps, guides VLM escape choices","Low light breaks VLM spatial maps; color does not","VLM memory diverges from human patterns, study shows"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00045,"raw_usage":{"total_tokens":2294,"prompt_tokens":1000,"completion_tokens":1294,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":616,"completion_tokens_details":{"reasoning_tokens":1230}},"tokens_in":616,"tokens_out":1294,"duration_ms":9749,"temperature":1.0,"reasoning_tokens":1230,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T00:27:15.735379+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Swap the semantic labels in the same four-room layout, relabeling the physical Conference Room as 'Storage Closet' and the Entrance as 'Conference Room' while keeping the inventory descriptions fixed to the visual contents. If models still choose the physically larger central room, the prior is visual or geometric; if they switch to the Entrance, the decision tracks the text labels, confirming the paper's claim. A complementary check is to compare the JSON map with a free-recall probe after the episode, without asking for JSON.","supporting_citations":[],"review_version":1}