{"id":"f0953870-fe9d-44e1-8b9f-f16a72a3e65e","arxiv_id":"2608.08904","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Action post-training of a VLM into a VLA produces a persistent depth-decodability floor and a specific late-layer cliff caused by interference from late MLP writes.","lead":"This paper finds that turning a vision-language model into a robot-control model degrades its internal ability to decode depth, with the largest loss in the final layers. The authors trace the late-layer damage to interference from MLP (feed-forward) blocks and show that deleting those writes partially restores depth readout.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The probing protocol never states that the VLM and VLA receive identical text prompts; if the prompts differ, the cross-model floor and cliff are confounded, undermining the headline claim.","rationale":"The reader's weakest assumption (probe comparability across models/layers) is real but partially defused by the paper's own wording: the headline claims are about decodability, and the ablation comparison in Sec. 4.2 is within-model, so cross-model format differences do not invalidate the causal localization. The prompt-comparability issue, by contrast, is an unacknowledged input confound. VLM/VLA comparisons in hidden states are only meaningful if the token sequences are identical at probe time; the manuscript never states this. Because attention mixes text and visual tokens, a different prompt changes the very representations being probed. This does not require any error in the ablation logic, but it changes what the paper can claim: the floor and cliff would be properties of the deployed VLA pipeline as a whole (weights plus prompt), not of action post-training per se. The test is cheap and decisive. I keep the verdict at CONDITIONAL because the central MLP-localization result may well survive; the paper needs to disclose and, if necessary, rerun with matched prompts before the cross-model headline is accepted.","tokens_in":11656,"tokens_out":14826,"duration_ms":147160,"concrete_test":"Check the exact text prompts used for Molmo2-ER and MolmoAct2-LIBERO in the layerwise probing (Sec. 4.1, Fig. 2). If they differ, rerun the full 2x36 DPT probing with an identical prompt for both models (e.g., the same LIBERO language instruction or the same fixed generic prompt) and recompute Table 2: the final-layer gap (0.246 d1) and the terminal slopes (+0.040 vs -0.123). If floor and cliff persist under identical prompts, the cross-model claims stand; if they shrink or invert, the findings are input-confounded. Also confirm that no action-expert tokens or different system prompts are included in the forward passes.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Sec. 4.1 specifies the image input (LIBERO frames, 256px, both cameras) but is silent on the text prompt tokens fed to Molmo2-ER versus MolmoAct2-LIBERO. Hidden states at visual-token positions are computed by attention over the full multimodal sequence (Eqs. 3-5), so any prompt difference (e.g., LIBERO task instruction for the VLA vs a generic caption prompt for the VLM) changes those hidden states independently of the action-post-training weight changes. This directly threatens the paper's central cross-model claims: the persistent floor (VLA below VLM at all 36 layers) and the late-layer inversion (VLM rises +0.040 over L28-L35 while VLA falls -0.123) could be input artifacts rather than effects of action post-training. The limitation section acknowledges probe-format confounds but never mentions prompt control. The Sec. 4.2 ablation localization is within-model and would survive a prompt mismatch, but the base-VLM control and the title's 'reduces' claim depend on input comparability. This is an unacknowledged, directly checkable confound, and it is more load-bearing than the probe-comparability concern the authors already disclaim in Sec. 3.3.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports a layerwise probing study comparing depth decodability in a weight-matched VLM/VLA pair, Molmo2-ER and MolmoAct2-LIBERO. Using a capacity-matched DPT head supervised by Depth-Anything-3, the authors measure depth accuracy (d1) at visual-token positions for all 36 decoder layers of both models. They find a persistent floor, with the VLA below the VLM at every layer, and a late-layer cliff, in which the base VLM rises by +0.040 over layers 28–35 while the VLA falls by −0.123. A symmetric ablation sweep over module types (MLP/attention), models (VLA/VLM), and twelve three-layer windows shows that ablating the final MLP window (L33–35) in the VLA raises final-layer d1 from 0.506 to 0.584, recovering the majority of the terminal drop, while attention ablations and the same intervention in the base VLM do not produce comparable recovery. Probing accumulated MLP writes shows that in the base VLM these deposits are more depth-decodable than the residual stream, whereas in the VLA they collapse over the final blocks. The paper concludes that action post-training reduces late-layer depth decodability by repurposing late MLP computation. The claims are empirical and the paper is explicitly self-aware about its single-seed and probe-capacity limitations.","tokens_in":11853,"tokens_out":5480,"duration_ms":56987,"significance":"If the findings hold, the paper provides a valuable mechanistic localization of VLA representation degradation: a specific module and layer window, with a subtractive intervention that recovers a majority of the terminal decodability drop. Strengths include the weight-matched model pair, the explicitly symmetric 2×2×12 ablation design, the use of publicly documented models and teacher, and the paper's unusually honest limitation statements. The ridge corroboration and the module-level decomposition are useful complements to the main probing result. However, two load-bearing issues prevent acceptance in the current form: the cross-model comparison does not state whether the two models receive identical text prompts, and the headline quantitative claims are made without any uncertainty quantification. The paper's contribution is important and potentially publishable, but these points need to be addressed first.","major_comments":[{"comment":"The probing protocol specifies image inputs, layer taps, and model identities, but never states that Molmo2-ER and MolmoAct2-LIBERO receive identical text prompts when the hidden states are collected. Because the hidden states at visual-token positions are computed by attention over the full multimodal sequence, any difference in text tokens (for example, a LIBERO task instruction for the VLA versus a generic caption for the base VLM) changes the layerwise d1 curves independently of the action-post-training weight changes. This directly threatens the headline floor and cliff, both of which are cross-model comparisons. The ablation localization in Sec. 4.2 is within-model and would survive, but the title's 'reduces' claim and the base-VLM control depend on input comparability. The limitation section acknowledges probe-format confounds but never mentions prompt control. Please report the exact prompt template used for each model and add a control in which both models receive identical prompts, or explicitly ablate text tokens so that the visual hidden states are computed under identical conditioning.","section":"Sec. 4.1; Eqs. (3)–(5)"},{"comment":"All layerwise d1 values and derived quantities (for example, the final-layer gap of 0.246, the terminal slopes +0.040/−0.123, and the ablation deltas in Table 3) come from a single seed with no reported variance. The mid-stack floor is as small as about 0.04 d1, and the base VLM's late 'recovery' is +0.040 over eight layers; without bootstrap intervals over rollouts or at least a small set of probe seeds, the reader cannot assess whether the floor and slope differences are reliable. The single-seed caveat in Sec. 4.2 and the Limitations section is an honest disclosure, but it is not a substitute for uncertainty quantification on the quantities that carry the central claims. Please add confidence intervals or error bars, or explicitly reframe the headline claims as single-seed observations whose quantitative magnitudes are not yet stable.","section":"Sec. 4.1, Table 2; Supplementary A"},{"comment":"The limitation stated in Sec. 3.3, that linear probes conflate representational content with representational format, applies equally to the nonlinear DPT-based d1 comparison used for the floor and cliff. Because a fresh DPT head is trained per layer and model, a difference in how depth is formatted (for example, a nonlinear or rotated code) would appear as a drop in decodability even if the same depth information were present. The paper's caveat in Sec. 3.3 is explicitly restricted to ridge scores, but the main d1 result has the same interpretation problem. A concrete control would be to probe late VLA states with a probe trained on VLM late states, or vice versa, or to train a small nonlinear readout and compare. In the absence of such a control, the conclusion in Sec. 4.3 that action post-training 'collapses depth decodability' should be stated as a claim about this probe family, not about geometric information content generally.","section":"Sec. 3.3, Sec. 4.3, Fig. 5"}],"minor_comments":[{"comment":"Figure 4 plots three curves (VLA-MLP, VLA-attention, VLM-MLP), but Table 3 reports a VLM-attention row; the figure does not show the VLM-attention ablation, so the reader cannot visually check the claim that this condition changes d1 by at most +0.010. Please add the VLM-attention curve to the figure or explicitly note in the caption that it is omitted for visual clarity.","section":"Fig. 4"},{"comment":"The statement that the sweep 'causally localizes the cliff to late VLA MLP computation' is slightly broader than the evidence: the cliff spans layers 28–35, while the ablation window covers only layers 33–35. The text distinguishes the terminal drop from the full cliff only implicitly. Please qualify the localization statement so that it refers to the terminal portion of the cliff, or add an ablation of the L28–32 window to establish that the earlier part of the cliff is also affected.","section":"Sec. 4.2"},{"comment":"The phrase 'massive-activation artifact' for the L16 MLP write spike in Fig. 7 is unexplained. Either provide a brief quantitative criterion (for example, a norm threshold or comparison with neighboring layers) or cite the massive-activation literature; as written, the label is not verifiable from the figure.","section":"Sec. 4.3, Fig. 7"}],"recommendation":"major_revision","confidential_remarks":"The prompt-control issue is the most serious: it is directly checkable and the paper currently does not rule it out. If the authors can confirm identical prompts for both models, or add an explicit control, and if they add uncertainty quantification for the headline numbers, the contribution is within the journal's scope and likely publishable. The single-seed reporting appears to be a robustness concern rather than an integrity issue. I see no concerns about novelty or citation patterns."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should read this one. The paper does something genuinely useful: it takes a weight-matched VLM/VLA pair and shows, with a symmetric full-stack ablation sweep, that the VLA's late-layer drop in depth decodability is specifically tied to its final MLP writes. That is a concrete, surgically addressable target for people building VLAs, and the design is careful—same probe capacity across all model-layer cells, rollout-split evaluation, attention ablations and base-model controls. Credit where due: the ablation sweep is hypothesis-neutral, and the module-level decomposition (depth riding the MLP deposits) explains why the late writes matter.\n\nThe soft spots are real but mostly fixable. The biggest one is something the stress-test caught and the paper never mentions: nowhere in Sec. 4.1 or the supplementary is it stated that Molmo2-ER and MolmoAct2-LIBERO received identical text prompts when the hidden states were extracted. Because visual-token states are computed by attention over the full multimodal sequence, a LIBERO task instruction for the VLA versus a generic caption for the VLM would change those states independently of the action-post-training weights. That directly threatens the persistent floor and the late-layer inversion—the two headline cross-model results. The within-model ablation localization survives a prompt mismatch, so the central causal claim about late MLP interference is less exposed, but the title claim 'action post-training reduces' depends on input parity. This is directly checkable and should have been reported.\n\nThe other issues are the ones the reader flagged: single-seed numbers without error bars (the authors acknowledge this and lean on the structured dissociation, which is fair), no code release, and the format-versus-content probe confound that is acknowledged but only partially mitigated by ridge corroboration. None of these are fatal; the paper is honest about them.\n\nVerdict: it deserves a serious referee. The prompt question needs to be answered, and variance estimates or multi-seed sweeps would firm up the headline numbers, but the core finding is interesting and the methodology is above the bar. I'd take it to reading group.","headline":"Solid mechanistic probing with a real late-MLP finding, but the paper must confirm prompt parity between the VLM and VLA before the cross-model floor and cliff claims are trustworthy.","tokens_in":12442,"tokens_out":3587,"would_cite":true,"duration_ms":35611,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Action post-training turns a vision-language model into a robot policy at the cost of depth decodability: the VLA loses depth at every layer and collapses in the final layers, and the collapse is caused by the last MLP writes.","keywords":["vision-language-action models","vision-language models","multimodal language models","spatial understanding","depth perception","representation degradation","layerwise probing","residual stream ablation"],"falsifier":"Train a matched-capacity nonlinear probe on the VLA's final-layer visual-token states; if depth decodability no longer collapses relative to the base VLM, the cliff is a linear-format artifact rather than a loss of represented depth.","tokens_in":11413,"feed_emoji":"🤖","tokens_out":8920,"duration_ms":78634,"temperature":0.7,"pith_summary":"This paper asks how much spatial understanding survives when a vision-language model (VLM) is post-trained into a vision-language-action model (VLA). Probing depth from every decoder layer of a weight-matched base VLM/VLA pair, it finds a persistent gap at every layer (the floor) and, on top of that, a late-layer collapse in the VLA that inverts the base model's final-layer recovery (the cliff). The paper causally localizes the cliff: deleting the final MLP writes raises terminal depth decodability from $d_1=0.506$ to $0.584$, recovering most of the $0.123$ drop, while attention ablations and the same MLP ablation in the base VLM do not. A module-level decomposition shows that the base VLM carries depth most readably in accumulated MLP writes, and action post-training collapses that pathway in the final blocks. The result matters because it separates a late-stage interference that can be undone by deletion from a persistent cross-layer gap in how robot policies inherit perception from language models.","feed_headline":"Turning a VLM into a robot policy costs it depth perception","feed_subtitle":"The terminal collapse traces to the last MLP blocks, and zeroing their writes recovers most of the drop.","key_machinery":"The load-bearing object is the residual-stream write decomposition: each decoder layer adds its attention and MLP outputs to the stream, $h^\\ell = x + \\sum_{i=0}^\\ell (a^i + m^i)$, so any single write can be removed and the stream re-read. The paper uses this decomposition to sweep three-layer ablation windows over the full stack for both modules and both models, and to probe the accumulated MLP deposits $\\sum_{i\\le\\ell} m^i$ separately from the stream $h^\\ell$. The probe is a capacity-matched Dense Prediction Transformer (DPT) head trained at every layer, supervised by a monocular depth teacher. The decomposition turns a descriptive layer-wise gap into a causal intervention: because the final MLP writes are additive terms, deleting them isolates whether their content interferes with depth decodability at the readout.","core_discovery":"MolmoAct2-LIBERO, a VLA produced by action post-training Molmo2-ER, decodes monocular depth worse than its weight-matched base VLM at every one of the 36 decoder layers, and the gap widens to $0.246$ in $d_1$ at the final layer: over layers 28-35 the base VLM's $d_1$ rises by $+0.040$ while the VLA's falls by $-0.123$, ending at $0.506$ versus $0.752$. A symmetric full-stack ablation sweep, testing MLP versus attention writes in both models across twelve three-layer windows, isolates the cause: zeroing the VLA's last MLP writes (layers 33-35) raises final-layer $d_1$ to $0.584$, recovering the majority of the terminal drop, whereas the same intervention in the base VLM and the attention ablations show no comparable effect. Probing accumulated MLP writes separately from the residual stream explains the dissociation: in the base VLM those deposits out-decode the stream at every layer, while in the VLA the deposits collapse over the final blocks from $d_1\\approx 0.68$ to $0.38$, dropping below the stream itself. The paper concludes that action post-training repurposes late MLP computation at the expense of geometric readout, and it is explicit that the ablation probes are single-seed, so the localization rests on the structured module-, layer-, and training-specific dissociation rather than on any isolated cell.","pith_inferences":["If the same late-MLP interference appears in other VLA families, a training-time objective that preserves depth decodability in the final MLP blocks could prevent the cliff without altering the action expert.","A nonlinear probe of matched capacity would separate two readings of the cliff: depth information destroyed versus depth information reformatted into a nonlinear code; the paper's ridge corroboration is within-model only and does not settle this across models.","The floor's presence at every layer hints that action post-training changes global representational formatting, not just late-layer content; if so, fixing the cliff alone would still leave a sizeable accuracy gap.","The same ablation sweep could be applied to other spatial primitives such as surface normals, object size, or affordance; if the late-window signature recurs, the cliff generalizes beyond depth."],"forward_implications":["The terminal cliff is not diffuse degradation: it is carried by the VLA's late MLP writes, since removing just the final three MLP windows restores most of the drop.","The base VLM's depth information lives primarily in accumulated MLP deposits rather than in the residual stream, so probing writes can expose geometric content that stream probes miss.","The floor and the cliff are distinct phenomena, so a remedy for one may not repair the other; late-layer write deletion is a proof-of-principle recovery for the cliff only.","Representation-level recovery from ablation does not imply closed-loop policy improvement; the paper makes no claim that the intervention improves robot behavior."],"supporting_citations":[{"why":"Supplies the weight-matched model pair (Molmo2-ER base VLM and MolmoAct2-LIBERO VLA) and the documented three-stage action post-training pipeline.","marker":"[9]"},{"why":"Establishes the drop-off-to-recovery layerwise template and the attention-knockout method used as a positive control and contrast.","marker":"[22]"},{"why":"Supplies the capacity-matched DPT probing protocol for geometric awareness that fixes probe capacity across layers and models.","marker":"[1]"},{"why":"Provides the Depth-Anything-3 monocular teacher whose predictions serve as pseudo-ground-truth depth targets.","marker":"[14]"},{"why":"Defines the Dense Prediction Transformer architecture used as the probe head.","marker":"[20]"},{"why":"Provides the residual-stream framework of additive attention and MLP writes that underlies the ablation and deposit-probing design.","marker":"[8]"},{"why":"Documents prior reports of representation degradation under action post-training that the paper extends from semantic to geometric targets and from description to causal localization.","marker":"[12]"},{"why":"Supplies the LIBERO observations and rollout split used for all probing and ablation experiments.","marker":"[15]"}],"fun_headline_variants":["Action post-training breaks a VLM's late-layer depth readout","How robot fine-tuning kills a VLM's late-layer depth perception","Robot training collapses a VLM's late-layer depth decoding","The hidden cost of robot policies: lost late-layer depth in VLMs","Turning a VLM into a VLA erases late-layer depth decoding"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes that a capacity-matched probe measures the same thing in both models, so the floor and cliff reflect what the hidden states represent rather than how linearly action post-training chose to format depth; if depth moved into a nonlinear code, the collapse could be a probe artifact.","fun_headline_variants_meta":{"raw":{"variants":["Action post-training breaks a VLM's late-layer depth readout","How robot fine-tuning kills a VLM's late-layer depth perception","Robot training collapses a VLM's late-layer depth decoding","The hidden cost of robot policies: lost late-layer depth in VLMs","Turning a VLM into a VLA erases late-layer depth decoding"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000638,"raw_usage":{"total_tokens":3321,"prompt_tokens":1079,"completion_tokens":2242,"prompt_tokens_details":{"cached_tokens":1024},"prompt_cache_hit_tokens":1024,"prompt_cache_miss_tokens":55,"completion_tokens_details":{"reasoning_tokens":2151}},"tokens_in":55,"tokens_out":2242,"duration_ms":292680,"temperature":1.0,"reasoning_tokens":2151,"cache_read_input_tokens":1024,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:21:25.140168+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train a matched-capacity nonlinear probe on the VLA's final-layer visual-token states; if depth decodability no longer collapses relative to the base VLM, the cliff is a linear-format artifact rather than a loss of represented depth.","supporting_citations":[{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Establishes the drop-off-to-recovery layerwise template and the attention-knockout method used as a positive control and contrast."},{"cited_title":"Transformer Circuits Thread (2021),https: //transformer-circuits.pub/2021/framework/index.html","cited_arxiv_id":null,"evidence_quote":"Provides the residual-stream framework of additive attention and MLP writes that underlies the ablation and deposit-probing design."}],"review_version":1}