{"id":"eb43de81-da1d-4982-a43d-f8670fdd4eef","arxiv_id":"2607.06925","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":4,"one_line_summary":"Goal-conditioned world models transcribe instructions instead of perceiving spatial relations when the instruction names the scored quantity, and removing the goal from the dynamics fixes it.","lead":"This paper shows that goal-conditioned world models can fake spatial reasoning by copying the instruction text rather than perceiving the scene, and proposes keeping the goal out of the dynamics as a fix. A smart generalist should read it because the failure mode and detection protocol apply broadly to any language-conditioned AI evaluated on a quantity named in its prompt.","discovery_kind":"unclear","skeptic_critique":{"model":"glm-5.2","headline":"The independence claim is near-tautological when the instruction fully names the answer; the informative regime (partial transcribability) is never tested.","rationale":"The reader correctly identified the dose-response as the weakest link and the synthetic-only load-bearing positive. I sharpen the mechanism: the independence test is near-tautological when the shortcut is complete (instruction fully names the answer), so it provides little information beyond 'shortcuts are used when available.' The more informative partial-transcription regime is untested. However, this does not change the verdict from CONDITIONAL. The practical prescription — remove the goal from the dynamics — follows from transcribability alone (part 1 of the claim), which is solidly demonstrated. The independence qualifier (part 2) is secondary: it motivates why 'better features won't help' but only in the full-transcription case the paper actually studies. The core finding (instruction transcription inflates grounding metrics, detectable via goal-withheld and counterfactual probes) is clearly demonstrated across three settings with validated controls. The remaining gaps (no code, no released model probed, control is a tie not a win, single-seed counterfactual) are honestly stated and appropriately temper the generality claim without undermining the central diagnostic contribution. CONDITIONAL remains the right call.","tokens_in":14579,"tokens_out":4217,"duration_ms":122420,"concrete_test":"Create a partially transcribable tabletop variant: the instruction names the relation but not the referents (e.g., 'move something to the left of something'), so the model needs both the instruction and the scene. Run the action-ablation dose-response (α: 1→0) on this variant and measure counterfactual cosine. If leakage increases as α decreases, predictor-competition holds in the partial-transcription regime and the independence claim fails; if it stays flat, the claim generalizes beyond the full-shortcut case.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim has two parts: (1) leakage occurs when the instruction names the scored quantity, and (2) leakage is 'essentially independent of how predictive the non-instruction inputs are.' Part (1) is well-supported across three settings. Part (2) rests on the dose-response (Table 3), but this test is near-tautological in the full-transcription setting: when the instruction directly names the answer, a complete shortcut exists, so the model uses it regardless of action quality — that is what shortcuts do. The synthetic dose-response (0.975→0.986 as α:1→0) confirms this but is unsurprising. The informative test would be a PARTIALLY transcribable instruction (instruction names the relation but not which objects are target/anchor), where the model must combine instruction and scene — there, degrading the scene or action could plausibly shift reliance toward the instruction, and predictor-competition would predict increased leakage. This regime is never tested. The only external independence test (Language-Table +direction, cosine 0.174→0.032) is acknowledged as low-signal (footnote 1: ~10× smaller per-step motion). Thus 'independent of non-instruction predictor strength' reduces to 'a complete shortcut is used when available,' which adds little beyond standard shortcut-learning intuition, and the abstract's claim that the protocol applies to 'any goal-conditioned world model' overreaches an evidence base of three 2D/symbolic settings with no partial-transcription test.","agreement_with_reader":"partial"},"referee_report":{"model":"glm-5.2","summary":"This paper identifies and characterizes a failure mode in goal-conditioned world models: when a language instruction directly names the scored relation (e.g., 'put X left of Y'), the dynamics predictor can transcribe the instruction into predicted anchors rather than perceiving the scene. The authors demonstrate this 'instruction leakage' on a controlled 2D tabletop, the external BabyAI benchmark, and Language-Table, using two clean controls: goal-withholding (collapsing readout accuracy from 0.90 to 0.27) and counterfactual goal substitution (94.5% false-instruction following). They further show that Language-Table, whose instructions name referents but not the scored direction, does not leak until the instruction is augmented to name the direction. The proposed fix—removing the goal from the dynamics and supervising the read path—recovers instruction-independent grounding (0.88, identical with and without the goal). The paper is methodologically careful, with validated positive controls, multi-seed replication for headline claims, and honest reporting of boundaries (the fix ties rather than beats the no-goal baseline on control; grounding collapses at maximal ambiguity).","tokens_in":15382,"tokens_out":1456,"duration_ms":125851,"significance":"The paper makes a valuable methodological contribution: it provides a falsifiable characterization of when instruction leakage occurs (transcribability), a validated detection protocol (goal-withheld and counterfactual probes with an engineered-leaky positive control at 0.97), and a simple architectural remedy. The cross-environment validation (tabletop, BabyAI, Language-Table) and the dose-response ablation (Table 3) are commendable. The honesty about the control result (a tie, not a win) and the ambiguity boundary (Fig. 4) strengthens the work. The characterization that leakage is governed by transcribability and is independent of non-instruction predictor strength is the central novel claim, and the three-setting evidence for the transcribability axis is solid. The independence sub-claim (Part 2 of the central claim) is more fragile, as detailed below.","major_comments":[{"comment":"§5.4, Table 3 and the abstract's central claim: The independence claim ('leakage is essentially independent of how predictive the non-instruction inputs are') rests on the action-ablation dose-response, but the load-bearing positive evidence is the synthetic regime (cosine 0.975→0.986 as α:1→0). This is near-tautological in the full-transcription setting: when the instruction directly names the answer, a complete shortcut exists, so the model uses it regardless of action quality—this is what shortcuts do by definition. The informative test would be a partially transcribable instruction (e.g., instruction names the relation but not which objects are target/anchor, forcing the model to combine instruction and scene). In that regime, degrading the action or scene could plausibly shift reliance toward the instruction, and predictor-competition would predict increased leakage. This regime is未","section":null},{"comment":"Abstract and §5.4: The claim that the protocol and remedy 'apply to any goal-conditioned world model whose instruction names the scored quantity' overreaches the evidence base. All three settings are 2D or symbolic gridworlds with discrete relations and short templated instructions. The authors acknowledge in §7(ii) that there is no 3D or real-world validation, and in §7(iii) that no released pretrained model was probed. The transcribability mechanism should generalize in principle, but the strength of the independence sub-claim and the universality of the remedy would be better supported by at least one higher-dimensional or continuous-relation setting. The authors should qualify the abstract's 'any' to match the settings tested.","section":null}],"minor_comments":[{"comment":"§5.2, Fig. 4: The across-ambiguity non-transfer result (a=2 model reads 0.86 at a=2 but falls to 0.26–0.33 at other ambiguities including easier a=0) is interesting but underexplored. A brief discussion of whether this is a memorization artifact or a fundamental limitation of per-ambiguity training would help the reader.","section":null},{"comment":"Table 1: The 'n/a' entries for ObjToken (object latents) make it hard to compare against the anchor-based models. A brief note on why the geometric readout cannot be applied (the slot convention issue is mentioned in §7(vi) but not at the table) would clarify whether this is a limitation of the metric or the design.","section":null},{"comment":"§3, 'PrismWM': The name appears without explanation of its etymology. A brief gloss would help.","section":null},{"comment":"§5.3: The control result is described as 'a tie, not a win' and the authors are commended for this honesty. However, the framing that 'goal-conditioning was the thing dragging GoalDyn down' could be read as slightly circular: GoalDyn is the authors' own architecture with goal-in-dynamics, so showing it underperforms NoGoal confirms the design choice but does not independently motivate it. A sentence acknowledging that this is a consistency check on their own design, not an external validation, would be fairer.","section":null},{"comment":"Fig. 2a: The y-axis label 'predicted-anchor accuracy (amb2)' could be misread as anchor localization accuracy rather than relation-readout accuracy. Clarifying that this is the geometric relation readout from predicted anchors would help.","section":null},{"comment":"§5.4, Language-Table: The counterfactual cosine for the +direction regime (0.174→0.032) is acknowledged as low-signal due to ~10× smaller per-step motion (footnote 1). The footnote's reasoning that the magnitude column 'refutes' the vanishing-motion floor is somewhat terse; a one-sentence explanation of why constant magnitude rules out the floor would help the reader.","section":null},{"comment":"References: Several citations are to 2026-dated arXiv preprints (Maes et al., Nam et al., Zhang et al., etc.). These should be verified for correctness and whether they have been published or updated since submission.","section":null}],"recommendation":"minor_revision","confidential_remarks":"The skeptic's concern about the independence claim being near-tautological in the full-transcription regime is partially valid: the dose-response is unsurprising when a complete shortcut exists. However, this does not undermine the paper's primary contribution (the transcribability characterization and detection protocol), which is well-supported. The independence sub-claim is a secondary part of the central claim and could be softened without losing the paper's value. The main actionable request is to either test a partial-transcription regime or qualify the independence and universality claims. This is a minor revision because the core contribution is sound and the issues are local to one sub-claim and the abstract's scope language."},"author_rebuttal":null,"desk_editor":{"model":"glm-5.2","letter":"The paper you should know about: Wang et al. identify a real evaluation confound in goal-conditioned world models — when the language instruction names the relation being scored, the predictor can just transcribe the instruction instead of perceiving the scene. A model hitting 0.90 relation-readout accuracy collapses to 0.27 when you withhold the goal, and follows a counterfactual (false) instruction 94.5% of the time. The fix is simple: keep the goal out of the dynamics, put it in the planner's cost, and supervise the read path. That's the whole story, and it's a good one.","headline":"Solid diagnostic finding with a clean fix; the independence sub-claim is the weak link","tokens_in":15317,"tokens_out":920,"would_cite":true,"duration_ms":30095,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"glm-5.2","headline":"When the instruction names the answer, world models copy, not perceive","keywords":[],"falsifier":"Feed a goal-conditioned world model a counterfactual instruction (one naming a relation different from the true scene) and measure whether the predicted anchors follow the false instruction or the true scene. If the model perceives rather than transcribes, the anchors should follow the true scene. If it transcribes, they follow the false instruction.","tokens_in":14613,"feed_emoji":"🔍","tokens_out":1991,"duration_ms":68164,"temperature":0.7,"pith_summary":"The paper asks when explicit reference anchors in a compact world model genuinely ground spatial relations versus when they merely appear to. The answer turns on a confound: if the language goal names the relation being scored, the model's predictor can copy the answer from the instruction rather than perceiving it from the scene. This instruction leakage inflates representation metrics (a goal-conditioned model reaches 0.90 relation-readout accuracy that collapses to 0.27 chance when the goal is withheld) and degrades control. The authors establish that leakage is governed by transcribability — whether the instruction names the scored quantity — and is essentially independent of how predictive the non-instruction inputs are, tested across a controlled tabletop environment, the BabyAI benchmark, and Language-Table. The remedy is architectural: keep the goal out of the dynamics, where it does not belong, and place it only in the planner's cost function, while supervising the perception (read) path. This recovers genuine, instruction-independent grounding at 0.88 accuracy, identical whether or not the goal is provided.","feed_headline":"World models copy the instruction instead of perceiving the scene","feed_subtitle":"A 0.90 grounding score collapses to chance when the goal is withheld. The fix: keep the goal out of the dynamics.","key_machinery":"The central mechanism is instruction leakage: a goal-conditioned dynamics predictor copies the relation named in the language instruction into its predicted anchor coordinates, bypassing scene perception entirely. The detection instrument is a pair of controls — goal-withheld (zero the goal tokens, recompute the readout) and counterfactual-goal (feed a goal naming a different relation, measure whether anchors follow the false instruction or the true scene). The fix is goal-free dynamics: the predictor never sees the goal, which enters only through the planner's cost function, combined with supervised supervision of the read (perception) path.","core_discovery":"The paper identifies and characterizes instruction leakage in goal-conditioned world models: when a language instruction directly names the spatial relation being evaluated, a goal-conditioned predictor achieves high relation-readout accuracy not by perceiving the scene but by transcribing the instruction into its predicted coordinates. Withholding the goal collapses accuracy from 0.90 to 0.27 (chance), and feeding a counterfactual instruction makes the model follow the false goal 94.5% of the time. The leakage is governed by transcribability — whether the instruction names the scored quantity — and does not depend on the predictive strength of the action or state inputs. Removing the goal 从","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Goal-conditioned world models leak instruction instead of perceiving scenes","World model grounds spatial relations by transcribing the instruction","Withholding the goal drops grounding accuracy from 0.90 to 0.27","Counterfactual instructions reveal world models copy rather than perceive","Goal-free dynamics recover instruction-independent grounding at 0.88"],"cache_read_input_tokens":0,"weakest_assumption_plain":"The claim that leakage is independent of non-instruction input strength rests primarily on a dose-response experiment in the authors' own synthetic tabletop environment, where per-step motion is large enough for the probe to be reliable. External benchmarks like Language-Table have roughly ten times smaller motion, making their probes low-signal, so the generalization from one synthetic environment to any goal-conditioned world model depends on that environment being structur","fun_headline_variants_meta":{"raw":{"variants":["Goal-conditioned world models leak instruction instead of perceiving scenes","World model grounds spatial relations by transcribing the instruction","Withholding the goal drops grounding accuracy from 0.90 to 0.27","Counterfactual instructions reveal world models copy rather than perceive","Goal-free dynamics recover instruction-independent grounding at 0.88"]},"model":"glm-5.2","effort":"low","cost_usd":0.0,"raw_usage":{"total_tokens":775,"prompt_tokens":688,"completion_tokens":87,"prompt_tokens_details":null},"tokens_in":688,"tokens_out":87,"duration_ms":44833,"temperature":1.0,"reasoning_tokens":null,"cache_read_input_tokens":0,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-09T22:47:43.181050+00:00","model_set":{"reader":"glm-5.2"},"falsifier":"Feed a goal-conditioned world model a counterfactual instruction (one naming a relation different from the true scene) and measure whether the predicted anchors follow the false instruction or the true scene. If the model perceives rather than transcribes, the anchors should follow the true scene. If it transcribes, they follow the false instruction.","supporting_citations":[],"review_version":1}