{"id":"c43d52f4-cf66-45fb-8d1b-b1ee0242259b","arxiv_id":"2602.23721","paper_version":2,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"StemVLA supervises a GPT-2-based VLA with predicted future 3D-geometry features (VGGT) and temporally aggregated history, reporting 86.0% on LIBERO-Long - but its CALVIN results and equations are placeholders.","lead":"This paper adds 3D scene-geometry prediction and 4D video-history memory to a vision-language-action robot model and reports strong success rates on simulated manipulation benchmarks. The preprint's headline CALVIN table, all method equations, and the code promised by its title are missing, so the central results cannot yet be checked.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"FSGWP ablation lacks a non-geometric control; the +19-point LIBERO-Long gain may be generic auxiliary supervision, not future 3D geometry.","rationale":"The reader's verdict is REJECT for incompleteness (missing equations, placeholders, no code) and data reliability (Table 3 arithmetic errors). My stress-test focuses on the scientific claim that would remain after the manuscript is completed: the FSGWP ablation does not distinguish geometric future knowledge from generic auxiliary supervision. The reader's weakest_assumption explicitly names this same confound ('the 19-point LIBERO-Long gain ... rather than by the extra training signal of any auxiliary loss'), so I agree with that identification. Because the manuscript as submitted fails independent completeness checks, my concern does not change the rejection; it reinforces it with a deeper issue that would require a new experiment to resolve. Even if the equations and CALVIN numbers were filled in, the current evidence would not support the mechanistic attribution unless the control experiment is added.","tokens_in":12900,"tokens_out":5975,"duration_ms":51909,"concrete_test":"Run a control ablation in which the FSGWP target is replaced by frozen CLIP features (or MAE features) of the future frame, keeping the spatial-geometric query, decoder, and L2 loss weight (λwk=0.1) unchanged, with the same LIBERO fine-tuning protocol and checkpoint selection. If the LIBERO-Long score remains near 86.0 instead of falling to 67.0, the improvement is not specific to VGGT 3D geometry.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central mechanistic claim—that explicit prediction of future 3D spatial geometry improves manipulation—rests on the FSGWP ablation in Table 4. Removing FSGWP drops LIBERO-Long from 86.0 to 67.0, Object 96.0 to 78.0, Spatial 96.0 to 76.5, and Goal 92.0 to 72.0. FSGWP bundles a learnable <spatial-geometric> query, a decoder, and an L2 loss against VGGT features of future frames (Eq. 6 is not printed in the manuscript). No control replaces the target with a non-geometric or non-future feature (e.g., current-frame VGGT features, or CLIP/MAE features of the future frame). The gain could therefore be caused by the regularizing effect of any auxiliary prediction loss, by temporal predictive training, or by the specific geometric content; the paper attributes it to the last without isolating it. If a generic target produces the same gain, the 'future 3D spatial-geometric world knowledge' explanation collapses, and the novelty claim is unsupported. Table 3 also contains arithmetic inconsistencies (CoT-VLA row average 69.0 vs computed 87.0), further weakening confidence in the reported magnitudes, but the missing control is the most load-bearing issue for the mechanism.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"StemVLA proposes a GPT-2-based vision-language-action model that augments 2D image and language inputs with (i) a 4D historical spatiotemporal representation obtained by feeding past frames through VGGT and a VideoFormer temporal aggregator, and (ii) a 3D future spatial-geometric world-knowledge prediction head (FSGWP) supervised by an L2 loss against VGGT features extracted from future frames. Actions are produced by a diffusion transformer conditioned on an action query. The paper claims state-of-the-art CALVIN ABC-D performance and strong LIBERO results, with ablations attributing large gains to the FSGWP module. The manuscript as submitted, however, is incomplete: the key equations are missing, the StemVLA CALVIN row and abstract numbers are placeholders, and the ablation design does not isolate geometric content from auxiliary supervision.","tokens_in":13072,"tokens_out":4312,"duration_ms":42734,"significance":"If fully substantiated, the core idea is timely and potentially useful: using a pretrained 3D reconstruction model as a target for latent future-geometry prediction, rather than pixel-level future-frame prediction, is a plausible way to inject spatial structure into a VLA. The paper also builds on externally validated components (VGGT, VideoFormer/VPP, DiT) and evaluates on standard benchmarks, so the empirical claims are in principle checkable rather than circular. The current manuscript, however, does not yet provide the evidence needed to assess these claims: the equation for the central FSGWP loss is absent, the authors' own CALVIN numbers are missing, and the ablation that carries the mechanistic story lacks the control necessary to distinguish geometric supervision from any auxiliary prediction loss.","major_comments":[{"comment":"Equations (1)–(7) are referenced in the text but none are actually printed. In particular, Eq. (6), the 3D Future Spatial-Geometric World Knowledge loss, and Eq. (7), the action diffusion loss, are central to the method, yet their functional forms, notation (o_t, H, z, feature dimensions), and loss weighting are undefined. Without these equations the FSGWP mechanism and the training objective are not checkable. The section must be rewritten with complete mathematical definitions.","section":"§3.1, §3.4, §3.5, Eqs. (1)–(7)"},{"comment":"The central CALVIN claim is unsupported by the manuscript as written. Table 2's StemVLA row consists entirely of 'xx.x' placeholders, and the Abstract and Introduction report the average sequence length as 'XXX' and improvement as 'XX.X% to XX.X%'. A state-of-the-art claim requires the actual numbers, the number of rollouts, and a description of the evaluation protocol. These placeholders are load-bearing and must be filled before the paper can be evaluated.","section":"§4.2, Table 2, Abstract"},{"comment":"The FSGWP ablation does not control for the type of auxiliary supervision. Removing FSGWP changes both the prediction target (VGGT features of future frames) and the presence of an auxiliary loss/decoder/query. A generic non-geometric or non-future target — e.g., current-frame VGGT features, future-frame CLIP/MAE features, or a random feature target — could plausibly produce the same regularization benefit. The +19-point LIBERO-Long gain is therefore not attributable to 'future 3D spatial-geometric world knowledge' without such a control. Please add at least one control target or substantially weaken the mechanistic interpretation.","section":"§4.3, Table 4, Q2"},{"comment":"The table contains arithmetic inconsistencies that undermine confidence in the reported magnitudes. The CoT-VLA row (Spatial 81.1, Object 87.5, Goal 91.6, Long 87.6) averages to 87.0, not the reported 69.0. The StemVLA row (96.0, 96.0, 92.0, 86.0) averages to 92.5, not the reported 92.0. These are not cosmetic issues: they affect the headline average-accuracy claim in the Abstract. Recompute and correct all averages, or state explicitly if the 'Average' column is computed over a different subset.","section":"§4.2, Table 3"},{"comment":"The manuscript is not in a reviewable state: the Conclusion contains 'XXXX success rate', the Abstract/Introduction contain 'XXX' placeholders, and the title promises 'Open-Source' but no code or checkpoints are provided anywhere in the text. Please complete all numeric entries, include a reproducibility statement/code link, and remove or resolve all placeholder tokens.","section":"§5, §6, and general manuscript readiness"}],"minor_comments":[{"comment":"The 'Pixel-wise Loss: 0.1' hyperparameter is not connected to any equation or described in the method. Clarify what it applies to and where it enters the total loss.","section":"§4.1, Table 1"},{"comment":"No standard deviations, number of seeds, or significance tests are reported for the LIBERO ablations. Given the magnitude of the claimed differences, at least multiple-seed results or error bars should be provided.","section":"§4.3"},{"comment":"The paper uses inconsistent names for the same ideas ('Spatial-Geometry' vs. 'Spatial-Geometric'; 'History Aggregator' vs. 'VideoFormer'), and some references are cited only by arXiv preprint numbers without version dates. Also, reference [20] is listed as the VPP paper but is also used as the source of VideoFormer; clarify the relationship.","section":"References and notation"},{"comment":"There are numerous typos and grammatical errors (e.g., 'integartes', 'benifical', 'secion', 'extend the world state n steps into the future' without defining n). A full language edit is needed.","section":"Throughout"}],"recommendation":"reject","confidential_remarks":"The manuscript appears to be an incomplete draft rather than a finished submission: core equations are absent, the authors' own benchmark numbers are placeholders, and reported averages are arithmetically inconsistent. These are not local issues that a minor revision can fix. I recommend rejection. If the authors provide a complete version with actual CALVIN results, fully specified equations, and the missing non-geometric control for the FSGWP ablation, the core idea may be worth reconsidering in a future submission."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is an unfinished draft whose central claims cannot be checked. The combination—GPT-2 MLLM with a learnable spatial query trained to match VGGT features of future frames, plus VideoFormer temporal aggregation, feeding a DiT action head—is genuinely new relative to the cited VLA work, and if the LIBERO numbers are real, the +19 point long-horizon gain from the FSGWP module would be worth knowing. But the manuscript as written does not support it.\n\nEquations (1)-(7) are missing, including the two losses (6) and (7) that the whole contribution rests on. Table 2 is all placeholders, and the abstract's CALVIN number is 'XXX'. For a paper titled 'open-source', there is no code or data link. Table 3 has arithmetic problems: CoT-VLA's row average is printed as 69.0 while its four scores average to about 87.0, and StemVLA's printed 92.0 average does not match its own four scores (92.5). That makes me distrust the transcribed baselines, and none of the gains have seed variance or error bars.\n\nThe biggest methodological hole is the FSGWP ablation: removing it drops LIBERO-Long from 86.0 to 67.0, but there is no control where the auxiliary target is replaced with, say, current-frame VGGT features or a generic future image feature. Without that, the 'future 3D geometry' story is just one possible explanation for what the extra loss does. Also, Section 4.3 credits the module with 'camera intrinsic/extrinsic parameters, depth information, point cloud data, and trajectory tracking'—none of which the method actually uses. That sentence should not be in the paper.\n\nCredit where due: the paper is honest about borrowing VGGT and VideoFormer, and the limitations section is reasonable about gripper constraints and motion quality. The design is coherent and the idea is testable. But the artifact is not close to a reviewable submission. As submitted, I would desk reject and ask for a completed version with the equations, CALVIN numbers, code, and a control ablation. If those come, I'd send it out.","headline":"A genuinely new recipe, but as submitted it is unverifiable: missing equations, placeholder CALVIN numbers, no code, and an ablation that cannot separate future-geometry supervision from any auxiliary loss.","tokens_in":13779,"tokens_out":5126,"would_cite":false,"duration_ms":50668,"reading_group":"no","serious_thinker":"yes","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"StemVLA claims that a robot policy improves on long-horizon tasks when the language model is trained to forecast future 3D scene geometry and to encode past frames as a 4D history, achieving state-of-the-art LIBERO and CALVIN results.","keywords":["vision-language-action","robot manipulation","3D spatial geometry","4D spatiotemporal representation","future prediction","long-horizon planning","diffusion policy","world model"],"falsifier":"Train StemVLA under identical conditions but replace the future-geometry target with a non-geometric future feature target—say, features from the same visual encoder with geometry information scrambled or a 2D future-frame feature target—while keeping the loss weight and setup fixed. If LIBERO-Long still improves from 67.0% to roughly 86.0%, the geometry interpretation fails; if it does not, the geometric content of the target is the active ingredient.","tokens_in":12595,"feed_emoji":"🤖","tokens_out":5646,"duration_ms":52484,"temperature":0.7,"pith_summary":"The paper is trying to establish that a vision-language-action model becomes a better long-horizon robot policy when it is explicitly trained to represent the 3D geometry of what comes next, not just the pixels of the present. StemVLA adds a future-oriented prediction head that regresses geometry features extracted from future frames, and a temporal aggregation module that fuses 3D features from past frames into a 4D historical representation. On the LIBERO benchmark the full model reports 92.0% average success and 86.0% on LIBERO-Long, with an ablation attributing a 19-point jump on LIBERO-Long (67.0% to 86.0%) to the future-geometry module alone. A sympathetic reader would care because this points to world-knowledge supervision in latent 3D space as an alternative to pixel-level video prediction for improving manipulation.","feed_headline":"Predicting future 3D geometry lifts robot task success by 19 points","feed_subtitle":"By forecasting scene geometry and folding past frames into a 4D history, StemVLA hits 92% average on LIBERO and tops CALVIN.","key_machinery":"The central object is the 3D Future Spatial-Geometric World Knowledge Predictor (FSGWP), a training-time head driven by a learned <spatial-geometric> query through the GPT-2-based language backbone; it predicts VGGT features of future frames and is trained with an L2 loss. The supporting object is the Historical Spatio-Temporal Encoder: VGGT, a pretrained visual-geometry transformer that produces depth-aware latent 3D features, plus VideoFormer, a temporal attention module that aggregates those features over time into a 4D historical representation. The <action> query then feeds a diffusion transformer that generates action sequences. The division of labor is concrete: the future-geometry he","core_discovery":"StemVLA's central claim is that explicit future spatial-geometric supervision is the missing ingredient in current VLA models. The model uses a pretrained visual-geometry transformer to extract latent 3D features from both historical and future observation frames; the future features act as regression targets, and the model is trained with an L2 loss to predict those future geometry features from current observations, language, and history. The same geometry encoder, followed by a temporal attention aggregator, builds a 4D historical representation of past motion. The paper reports that this combination yields state-of-the-art average completed sequence length on CALVIN ABC-D and 92.0% avera","pith_inferences":["The mechanism suggests a broader design principle: latent 3D geometry is a better auxiliary training signal than raw pixels. A testable extension is to compare the same loss against future-frame feature targets from a 2D encoder; if the geometry target wins, the 3D content is doing the work.","One implicit consequence is that future-geometry supervision could be applied to non-language policies or even perception-only models, decoupling the representation benefit from the language backbone.","The paper leaves open how far ahead the 'n steps' horizon should reach; the +19-point result hints that horizon length itself is a knob worth sweeping, especially on longer benchmarks.","The design also implies that a purely RGB-based VLA can acquire depth-like spatial awareness without ever seeing depth labels at training time."],"forward_implications":["If the future-geometry mechanism is what drives the gains, then other VLA architectures can adopt the same training-time regression target without changing their inference-time policy.","The largest ablation jump appears on LIBERO-Long, suggesting explicit 3D forecasting matters most as task horizon grows.","Because the prediction heads are dropped at inference, the model gets the benefit of world-knowledge supervision at no extra deployment cost.","The 4D history encoder offers a reusable way to inject temporal structure into language-model-based policies beyond frame-stacking."],"fun_headline_variants":["Future 3D prediction boosts robot success to 92% on LIBERO","StemVLA forecasts 3D future and 4D history, hits 92% on LIBERO","Open-source model with future 3D knowledge achieves 92% on LIBERO","Predicting 3D scene geometry lifts robot success to 92% on LIBERO"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The claim rests on the assumption that matching an L2 prediction target to VGGT features of future frames teaches the model genuinely geometric future knowledge, and that the 19-point LIBERO-Long gain comes from that geometry rather than from any auxiliary prediction signal.","fun_headline_variants_meta":{"raw":{"variants":["Future 3D prediction boosts robot success to 92% on LIBERO","StemVLA forecasts 3D future and 4D history, hits 92% on LIBERO","Open-source model with future 3D knowledge achieves 92% on LIBERO","Predicting 3D scene geometry lifts robot success to 92% on LIBERO"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001004,"raw_usage":{"total_tokens":4120,"prompt_tokens":815,"completion_tokens":3305,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":559,"completion_tokens_details":{"reasoning_tokens":3212}},"tokens_in":559,"tokens_out":3305,"duration_ms":20065,"temperature":1.0,"reasoning_tokens":3212,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T20:12:29.940059+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train StemVLA under identical conditions but replace the future-geometry target with a non-geometric future feature target—say, features from the same visual encoder with geometry information scrambled or a 2D future-frame feature target—while keeping the loss weight and setup fixed. If LIBERO-Long still improves from 67.0% to roughly 86.0%, the geometry interpretation fails; if it does not, the geometric content of the target is the active ingredient.","supporting_citations":[],"review_version":1}