{"id":"0c9c76da-6b09-40d5-8258-4aec70ccfac2","arxiv_id":"2607.23797","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"Persistent map memory schedules budgeted re-perception better than memoryless VLM priors, with gain equal to Var(√λ), and language-conditioned VLMM needs both open-vocabulary relevance and per-instance dynamics.","lead":"A robot map that remembers how objects move is better at deciding what to look at next than at deciding how to walk around a room. Under a limited sensing budget, that memory beats on-demand vision-language guesses and helps most when a spoken instruction names what matters.","discovery_kind":"new_application","skeptic_critique":{"model":"moonshotai/kimi-k3","headline":"The distinctive “motion channel beats a timestamp” claim rests on a small +2.5% margin in §III-F that may not be cluster-robust across only 8 queries/scenes.","rationale":"I partially agree with the reader: the sim-to-real gap in the change process, unknown natural h, and the ~0.3 observation-reliability crossover are real external-validity conditions, and the paper states them plainly. I would not change the CONDITIONAL verdict. The additional stress point is internal/statistical rather than a contradiction: the most novel and specific part of the strongest claim is not the Var(√λ) law—which is largely algebraic under the stated surrogate and is caveated—or the +21–26% CLIP result—which the authors decompose into weak-prior correction plus a ~5% heterogeneity slope—but the +2.5% victory over relevance-weighted recency. That is the evidence that the motion channel adds value beyond a last-seen timestamp; if it is not robust to query/scene clustering, the central “vision–language–motion together” conclusion becomes much thinner even if the memoryless-prior result stands. The proposed clustered/permutation re-analysis is cheap relative to the claim and would either harden the distinctive contribution or clarify that the robust claim is persistence rather than rate modelling. Keep correctness risk medium and confidence moderate; acceptance should require both real longitudinal/reliability validation and cluster-robust language-conditioned separation.","tokens_in":13136,"tokens_out":6115,"duration_ms":272015,"concrete_test":"Re-run §III-F with a pre-registered held-out suite of ≥60 instructions across new AI2-THOR scenes; score VLMM vs relevance-oldest with a mixed-effects model or cluster bootstrap by query and scene (objects only within clusters), plus null permutations of s_i(q) and of λ_i labels. Report influence per query. Accept the “both channels” claim only if the cluster-robust 95% CI for the gap excludes 0 by a preset margin (e.g. >0.5%) and the relevance/λ permutations collapse the gap; otherwise downgrade to the weaker observation-vs-prior claim.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The broad result—using past observations at all beats a memoryless category/VLM prior—is well supported and honestly bounded (floor subtraction, Whittle reference, noise/reliability reversals). The sharper load-bearing claim is the paper’s distinctive one: in language-conditioned re-perception, VLMM (open-vocabulary relevance × observed dynamics) beats relevance-weighted oldest-first by +2.5% (t=7.4), so the rate channel is “more than a timestamp.” That margin is modest and comes from an online arena with 159 objects but only 8 queries; objects are nested within scenes/queries, and CLIP relevance scores are shared/similar within categories, so an object-level t-statistic can overstate effective sample size. Because relevance and λ are made nearly independent, the language channel is clearly necessary, but the necessity of the dynamics channel depends on this +2.5% being stable across query choice and correlation structure. The paper itself flags the setting-dependence (§III-G fetch: recency ties/beats memory; Limitations v), so if the language-conditioned gap is driven by a few queries or by clustered measurements, the defensible claim shrinks to “observation-based scheduling beats on-demand priors,” not “both VLMM channels are jointly necessary.” This is an inferential robustness issue, separate from the already-noted sim-to-real change-process concern.","agreement_with_reader":"partial"},"referee_report":{"model":"moonshotai/kimi-k3","summary":"The manuscript contrasts two uses of a persistent Vision–Language–Motion Map (VLMM). For spatial navigation, behavior-aware path costs improve a planning-time objective but provide little closed-loop benefit. The main contribution reframes budgeted re-perception as attention: objects are assumed to change as point processes, a √-law allocates observation frequency, and a Cauchy–Schwarz argument predicts that the benefit of instance-level memory grows with root-volatility heterogeneity. Experiments in AI2-THOR with simulated change processes compare category priors, held-out change histories, oldest-first and Thompson schedulers, oracles, a Whittle-index policy, and a real CLIP prior. Memory generally performs best as heterogeneity increases, with robustness, reliability, fetch-task, and language-conditioned studies. The distinctive final claim is that language-conditioned VLMM scheduling beats relevance-weighted recency by +2.5% and an on-demand VLM by +8.9%, implying that both language grounding and per-instance dynamics are needed.","tokens_in":13475,"tokens_out":9775,"duration_ms":208174,"significance":"If the claims are appropriately qualified, the paper makes a useful conceptual contribution: persistent behavior maps may matter less for global path shaping than for deciding what to re-perceive. The closed-loop negative navigation result is valuable, and the attention study has several methodological strengths: transparent √-law and Cauchy–Schwarz arguments, a held-out rate-estimation protocol, explicit subtraction of the h=0 schedule floor, oracle and Whittle-index references, competitive age and Thompson baselines, process and observation-noise variants, real CLIP features, reliability reversal tests, and a falsifiable heterogeneity-scaling prediction. The work is therefore potentially influential for lifelong mapping and active perception, although its practical magnitude remains conditional on natural scene heterogeneity and observation reliability.","major_comments":[{"comment":"The linear objective relies on λ_iτ_i≪1, but the default λ_max=0.12 and K=0.06N imply a mean revisit period near N/K≈16.7 and λτ as large as about 2. Prop. 1 also imposes only Σf_i=K, omitting the physical bounds 0≤f_i≤1 and the mapping to discrete top-K observations. The √-law therefore need not be a feasible optimum in the reported regime. Please quantify the surrogate error, impose or rule out binding box constraints, and compare with a schedule derived from the exact g(λτ).","section":"§II, Eq. (1) and Proposition 1"},{"comment":"Eq. (2) is derived for a homogeneous/global-mean memoryless baseline, whereas the experiments use category-varying priors ρ_i. At h=0, λ_i=λ_maxρ_i, so Var(√λ) is nonzero (reported as about 0.0009), yet the true memory-attributable gain is zero; Fig. 5's nonzero x-intercept makes the same point. Thus the headline claim that the gain “exactly equals Var(√λ)” is not the identity for the actual category prior. Please derive the Cauchy–Schwarz defect for λ̂=ρ_category, or use residual/within-category heterogeneity unexplained by that prior, and revise the abstract accordingly.","section":"§II, Corollary 1 and Eq. (2); §III-D, Fig. 5"},{"comment":"The paper's most distinctive claim—that the motion channel is more than a timestamp—rests on a +2.5% margin over relevance-weighted oldest-first, with t=7.4 but only 8 queries and 159 objects. The sampling unit and degrees of freedom are not stated, and objects are clustered by scene/category while sharing CLIP relevance and change histories; an object-level t-test may overstate effective sample size. Please report query- and scene-clustered confidence intervals or bootstrap tests, the per-query effects, and leave-one-query/scene-out results. Also specify how possibly negative CLIP cosine scores are transformed, since √(s_iλ̂_i) and s_i·age_i require nonnegative importance.","section":"§III-F, Fig. 7"},{"comment":"The central effect is demonstrated under generated λ_i=λ_max[(1−h)ρ_i+hu_i], while the natural value and dependence structure of h are unmeasured. The bursty, correlated, and diurnal variants are useful, but they remain mean-matched synthetic processes. Spatially or task-correlated shocks and stronger within-category homogeneity could materially shrink per-instance memory value. Please either estimate heterogeneity from real repeated scans/longitudinal data or add a hierarchical/correlated-shock sensitivity study and report gain against residual heterogeneity. At minimum, scope the conclusions explicitly as a simulation-established mechanism rather than a measured deployment effect.","section":"§III Setup, §III-C, and §IV(i–ii)"},{"comment":"The +21–26% result uses a zero-shot CLIP movability prior whose correlation with ground truth is only r=0.40. The manuscript appropriately notes that most of the base gain corrects this weak appearance prior, that a clean category prior beats memory below h=0.5, and that the base advantage shrinks as prior quality improves. The broader statements that an “on-demand VLM” is a poor scheduler are therefore too general. Please name the tested CLIP baseline in the abstract/conclusions, separate prior-correction gain from the heterogeneity slope, and ideally evaluate a stronger contextual VLM before making a general VLM claim.","section":"§III-E, Table IV and Fig. 6; Abstract"}],"minor_comments":[{"comment":"The symbol h denotes both per-instance heterogeneity and, in Eq. (3), the age-dependent holding cost h(τ). Rename the latter, e.g. c(τ), and avoid overloaded notation.","section":"§III-C, Eq. (3)"},{"comment":"Fig. 1 labels the budget B, while the text and Proposition 1 use K. Please unify the notation.","section":"Fig. 1 and §II"},{"comment":"The exact query strings, number of independent runs/seeds, and unit underlying t=7.4 should be reported. Error bars are needed on the bars, and all five bar values should be legible.","section":"§III-F, Fig. 7"},{"comment":"The numerical definitions of “importance skew” 0/0.5/1/2, the log-normal importance parameters, and uncertainty intervals for each table entry are missing. These details are needed to reproduce the sensitivity conclusions.","section":"Table III"},{"comment":"Please give the exact movable/fixed prompt pair, score calibration, and definition of the correlation r. This will help distinguish CLIP's movability estimate from a broader assessment of modern VLM capability.","section":"§III-E"},{"comment":"The reliability reversal is important, but only four mean reliability levels are shown. Please identify the detector, explain how measured detectability is scaled, report a finer curve with uncertainty, and state whether moving to a usable viewpoint consumes the same perception budget.","section":"§III-H, Table V"},{"comment":"The 35% planning-time objective, 4% closed-loop reduction, 14% travel penalty, and replan counts need formal metric definitions and preferably a table. This negative result is valuable and deserves the same reporting precision as the attention experiments.","section":"§III-I"}],"recommendation":"major_revision","confidential_remarks":"The manuscript builds directly on the same author’s VLMM preprint [1] as its map substrate. This is disclosed, but the editor may wish to verify that [1] is accessible and that the contribution boundary is clear. The principal editorial risk is breadth of claims relative to the simulated longitudinal dynamics; this appears correctable through the analyses and claim narrowing requested above."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"The useful punchline is the negative plus the reframe. Behavior-aware path costs look strong at planning time (~35%) and nearly vanish in closed loop (~4%), while the same map’s change history (or even plain recency) is a good scheduler for budgeted re-perception and beats a memoryless VLM prior. That distinction is worth having on the record.\n\nWhat is actually new is not the √·-law or Whittle index—they cite the web-crawler and restless-bandit literature correctly—but the clean spatial-vs-attention split, the held-out protocol that isolates memory-attributable gain, and the empirical match of that gain to Var(√λ) (R²=0.99). Prop. 1–2 and Cor. 1 are standard convex/Cauchy–Schwarz arguments and they line up with the plots. They do the right controls: floor subtraction at h=0, Whittle reference that removes the floor, oldest-first/Thompson baselines, process variants, noise and real-detectability sweeps, and a real CLIP prior. The language-conditioned ablation is the distinctive claim and is honestly bounded in the limitations.\n\nSoft spots in proportion. The load-bearing change process is still parametric and simulated; natural h is unknown, and they say so. Memory reverses below ~0.3 observation reliability, which is a real deployment condition, not a footnote. The sharper “motion channel beats a timestamp” result is only +2.5% (t=7.4) over relevance-weighted oldest-first on 8 queries/159 objects; objects nest in scenes/queries and the paper itself notes that on the generic fetch task recency ties or wins. I would not hang the joint-necessity slogan on that margin alone. The broad claim—using past observations beats an on-demand category prior—is on firmer ground.\n\nNo code release. Citations look appropriate; self-cite to VLMM is substrate, not circularity. For people working on lifelong open-vocab maps and active perception this is worth a careful read. I would send it to referees: the framing and the negative are clear enough, the math is sound, and the sim-to-real gaps are stated rather than hidden. Engage, but treat the +2.5% as provisional.","headline":"Solid reframe: behavior maps pay off as budgeted attention, not path costs; math and held-out evidence are clean, but the distinctive +2.5% motion-over-timestamp claim is thin and dynamics remain simulated.","tokens_in":14703,"tokens_out":571,"would_cite":true,"duration_ms":10943,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"A behavior map earns its keep by telling a robot what to re-observe under a budget, not how to walk around a room.","keywords":["vision-language-motion maps","budgeted re-perception","attention scheduling","open-vocabulary mapping","restless bandits","age of information","robot memory","language-conditioned perception"],"falsifier":"Run the same held-out and language-conditioned schedules on a real robot with longitudinal, naturally occurring object moves and real detection success rates; if memory no longer beats recency and the on-demand prior once natural heterogeneity and viewpoint-coupled observation replace the simulated λ_i model, the central claim fails.","tokens_in":14501,"feed_emoji":"🤖","tokens_out":961,"duration_ms":25464,"temperature":0.7,"pith_summary":"A robot with a persistent map that records what moves faces two planning jobs. Shaping global paths around possible motion looks helpful on paper but nearly vanishes once the robot is actually driving and can replan reactively; an on-demand vision–language query does about as well. The job where memory matters is attention under a limited perception budget: which map entries to re-check so the representation stays fresh. History of actual change (or even simple recency) produces the best held-out schedule and matches an oracle, while a memoryless category prior is a poor scheduler. The gain equals the variance of root change-rates, concentrates on the objects that matter, and is largest when language names what to track—so open-vocabulary grounding and per-instance dynamics are both required.","feed_headline":"Map memory tells robots what to watch, not where to walk","feed_subtitle":"Under a perception budget, change history beats on-demand vision–language priors; path-shaping gains mostly vanish.","key_machinery":"The √·-law attention schedule (re-check frequency proportional to √(w_i λ_i)) and its Cauchy–Schwarz consequence that the freshness gap versus a uniform/memoryless prior equals Var(√λ), the variance of root-volatility; a Whittle-index restless-bandit policy serves as the discrete reference that removes schedule artifacts.","core_discovery":"Under a limited perception budget, a persistent map’s memory of per-instance change yields the best re-perception schedule—matching an oracle and beating a memoryless on-demand VLM prior—with the memory-attributable gain equal to Var(√λ). In the language-conditioned setting, open-vocabulary relevance times observed dynamics beats even relevance-weighted recency and an on-demand VLM; neither language nor dynamics alone suffices. Spatial path-shaping benefit largely disappears in closed-loop execution.","pith_inferences":["The same Var(√λ) logic suggests lifelong mapping systems should allocate mapping compute proportional to estimated root-volatility rather than uniform coverage sweeps.","Household and warehouse deployments with high instance-to-instance usage differences (same chair type used very differently) are the natural first places the claimed gain should appear at scale.","A practical system could gate the rate channel on measured observation reliability and importance skew, defaulting to cheapest recency otherwise.","The closed-loop path-shaping negative implies many ‘dynamic costmap’ papers may be over-claiming once reactive local control is admitted."],"forward_implications":["Map designers should treat behavior annotation primarily as an attention/resource-allocation signal, not as a global path-cost field.","When perception budget is scarce, storing per-instance change history (or at least last-seen times) is worth more than querying a category-level VLM on demand.","Language-conditioned tasks need both open-vocabulary grounding and a dynamics channel; a static open-vocabulary map plus recency is not enough.","The quantitative value of memory is predicted by measurable root-volatility variance, so scenes can be scored for whether a persistent map will pay off.","If observation reliability falls much below ~0.3 mean success, policies should fall back to the category prior rather than trust corrupted history."],"fun_headline_variants":["Map memory picks what to re-observe, not how to walk","Change history beats on-demand VLM under perception budgets","Language-conditioned memory tells robots what merits attention","Persistent map matches oracle re-perception; path gains vanish","Memory gain equals Var(√λ); focuses budget on volatile objects"],"cache_read_input_tokens":128,"weakest_assumption_plain":"Object change is assumed to follow independent parametric processes whose per-instance heterogeneity can be dialed by a single knob; real long-term scene dynamics and natural heterogeneity are not measured.","fun_headline_variants_meta":{"raw":{"variants":["Map memory picks what to re-observe, not how to walk","Change history beats on-demand VLM under perception budgets","Language-conditioned memory tells robots what merits attention","Persistent map matches oracle re-perception; path gains vanish","Memory gain equals Var(√λ); focuses budget on volatile objects"]},"model":"grok-4.5","effort":"low","cost_usd":0.004517,"raw_usage":{"total_tokens":1478,"prompt_tokens":979,"num_sources_used":0,"completion_tokens":85,"cost_in_usd_ticks":45168000,"prompt_tokens_details":{"text_tokens":979,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":414,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":979,"tokens_out":85,"duration_ms":10065,"temperature":1.0,"reasoning_tokens":414,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-30T12:04:54.656962+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"Run the same held-out and language-conditioned schedules on a real robot with longitudinal, naturally occurring object moves and real detection success rates; if memory no longer beats recency and the on-demand prior once natural heterogeneity and viewpoint-coupled observation replace the simulated λ_i model, the central claim fails.","supporting_citations":[],"review_version":1}