{"id":"b12c1bfb-16aa-4625-96a2-51875298f0e5","arxiv_id":"2607.26657","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"Enfold folds the internal computation of a video world generator into a current-only representation, enabling competitive robot control without executing the generator at deployment.","lead":"Enfold trains a robot policy to predict the internal states a video-generation model would produce while imagining a future, so at run time the robot can act without running the generator. The method matches or slightly exceeds prior world-action models on LIBERO and RoboTwin while cutting action-prediction latency by 3.7–10.1×.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"G2R target's causal role is untested: no control with an alternative future-predictive target under the same detached protocol, and no reported G2R prediction accuracy.","rationale":"The reader's weakest assumption focused on the post-hoc selection of supervision layers, which is a valid robustness concern but does not directly threaten the mechanism. My concern is more load-bearing: the paper does not demonstrate that the generator-state target specifically is necessary, and it does not quantify how well the encoder learns the target. A proper control with an alternative future-predictive representation (e.g., frozen DINOv3 future features) and measurement of G2R prediction accuracy would settle whether the generator's internal computation is actually the source of the gains. Since this is an untested assumption rather than a demonstrated flaw, the conditional verdict remains appropriate.","tokens_in":22877,"tokens_out":12164,"duration_ms":188218,"concrete_test":"Add an ablation 'Future-DINO' that replaces r^G_t in Eq. 4 with future-frame DINOv3 ViT-H features under the same timestep-conditioned head, detached task readout, and training budget; also report the held-out cosine similarity between \\hat{r}^G_t and r^G_t for Enfold. If Future-DINO matches Enfold on LIBERO (within 0.5 points) or Enfold's G2R similarity is low (< 0.6), the causal role of the generator's internal states is not established.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the reusable asset is the world generator's internal computation, transferred via G2R. The paper's only evidence that generator states specifically (rather than any strong future-predictive target) drive the gains is Table 4, which compares against future pixels and action-only. Future pixels are a much more difficult target (dense reconstruction) and action-only training changes the gradient protocol (task gradients shape the encoder), so neither is a controlled substitute for the generator-state target. No baseline predicts future features from a frozen encoder (e.g., DINOv3 or V-JEPA) under the same architecture, same loss, same stop-gradient task head, and same compute. Moreover, the paper never reports how well the current-only encoder actually predicts the multi-level generator target (G2R loss / cosine similarity on held-out clips). If the G2R target is only weakly predictable, the observed performance could arise from the strong base features plus the auxiliary objective, not from internalizing the generator's computation. This is a load-bearing gap because it attacks the mechanism, not just the layer selection.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Enfold, a method for transferring the internal computation of a video world generator into a current-only representation. During training, multi-level generator states obtained from a teacher-forced pass over the observed future supervise a current-only predictive encoder (G2R), the learned representation conditions future generation (R2G), and task heads read the representation with stop-gradient. At deployment, the generator is not executed for action prediction. The authors report strong control results on LIBERO (97.8% average success), RoboTwin2.0 (91.77%), and real bimanual tasks, with 3.7x lower latency than Fast-WAM and 10.1x for a TensorRT-accelerated variant. Additional analyses probe layer selection, future-prediction improvements from R2G, representation statistics, and human-intervention behavior.","tokens_in":23149,"tokens_out":3848,"duration_ms":62791,"significance":"If the central claim holds, the paper makes a useful conceptual and practical contribution: world models can serve as training-time sources of structured predictive supervision rather than as mandatory inference-time simulators. The latency gains are concrete, and the detached task-readout design cleanly separates representation learning from policy optimization. The paper is commendably transparent about several limitations, including the small layer-selection probe, the qualitative nature of the intervention evidence, and the need for uncertainty estimates on close benchmark margins. However, the mechanism claim — that the specific generator-state target, rather than future-predictive supervision in general, drives the gains — is not yet isolated by the reported experiments.","major_comments":[{"comment":"No uncertainty intervals are reported for any success rate, and several headline margins are very small: Enfold is 0.2 points above Fast-WAM on LIBERO average and 0.19 points above on RoboTwin2.0. The text itself states that these margins are too small to support a ranking claim without uncertainty estimates. Please report standard errors, confidence intervals, or multiple-seed evaluations, and adjust the strength of the comparative claims accordingly. This is essential for the central efficiency-accuracy trade-off claim.","section":"Section 5.2, Tables 1-2"},{"comment":"The G2R ablation does not isolate the causal role of generator states as the supervision target. Future-pixel prediction is a much harder reconstruction objective, and action-only training changes the gradient protocol by letting task gradients shape the encoder. Neither is a controlled substitute for a future-predictive feature target. The paper should include a baseline in which the current-only encoder predicts future features from a frozen visual encoder (e.g., DINOv3 or V-JEPA) under the same architecture, loss, stop-gradient task head, and compute. In addition, the paper never reports how well the current-only encoder actually predicts the multi-level generator target on held-out clips. Reporting G2R loss or cosine similarity would directly substantiate the claim that the generator's internal computation is 'internalized' rather than only weakly correlated with a strong auxiliary o","section":"Section 5.4, Table 4; Section 4.2, Eqs. (3)-(4)"},{"comment":"The supervision layers L={7,15,23,27} are a load-bearing design choice, but the selection is post hoc on a 20-video probe. The appendix explicitly cautions that this probe is insufficient to establish a precise layer ranking, and Figure 2 shows no universally best layer. Because the G2R target and all downstream results depend on this fixed set, the paper should provide additional evidence that the choice is stable: e.g., a sensitivity analysis over alternative layer sets, a cross-task layer-selection experiment, or a validation-set criterion that does not reuse the final test set. Without this, the target specification remains an unexplained empirical choice.","section":"Section 3, Appendix C"},{"comment":"The human-intervention evidence is qualitative only, and the appendix states that no matched frozen-representation or no-intervention rollout is reported. The claim that the method 'adapts' rather than replays is central to the paper's narrative, but the current evidence is a few hand-picked rollouts. Please either add quantitative intervention experiments with multiple seeds and a recovery-rate metric, or clearly demote this claim to a qualitative illustration and remove it from the abstract-level conclusions about counterfactual consistency.","section":"Section 5.3, Appendix B.5"}],"minor_comments":[{"comment":"The symbol LN is used without definition. Presumably layer normalization; please define it at first use.","section":"Section 4.2, Eq. (3)"},{"comment":"The first row label 'F uture pixels' contains a spacing artifact; should be 'Future pixels'. Similar spacing artifacts appear in Section 5.5 headings and elsewhere, e.g., 'F rom', 'T oken'. A careful proofread is needed.","section":"Section 5.4, Table 4"},{"comment":"Several table cells appear to be improperly typeset (e.g., '92.293.3' and '86.183.3'). Please check all numerical entries for formatting and column alignment.","section":"Table 3 and Table 7"},{"comment":"The effective-rank comparison uses different sample sizes (20 clips in Appendix C, 100 clips in Figure 6). The text says raw values are not compared across sets, which is appropriate, but the caption of Figure 6 should state this explicitly to avoid reader confusion.","section":"Appendix B.4"},{"comment":"The information-theoretic statement I(Y;Z|X)=0 is correct for deterministic U_phi but may be surprising to readers because Z is a function of X. Please add one sentence clarifying that this is a statement about the learned encoder at a fixed time, and that R2G gains reflect reorganization, not additional information about Y.","section":"Appendix A, Eq. (14)"}],"recommendation":"major_revision","confidential_remarks":"This is a well-written and honest paper with a fresh idea and strong empirical breadth. The main technical gap is that the mechanism claim is not yet isolated from the general benefits of future-feature predictive supervision, and the benchmark margins need uncertainty quantification. These are fixable with additional experiments. I recommend major revision rather than rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know: this is a real result, not a framework paper. The authors show you can take a pretrained video generator, use its internal states during training as supervision for a current-only encoder, and then drop the generator from the action path while keeping competitive success rates on LIBERO, RoboTwin2.0, and a real robot. The latency numbers (134 ms, 49 ms with TensorRT) are measured on the same protocol and are credible.\n\nThe paper does several things well. The ablations in Table 4 separate multi-level generator states from single-level, future pixels, and action-only; the multi-level target wins, and the gain concentrates on goal/long-horizon suites. The R2G direction is a nice functional test: conditioning the generator on the learned representation improves future video prediction (PSNR, SSIM, and LPIPS all move in the same direction). The representation analysis shows the encoder is not just a compressed copy of a teacher layer—it has higher effective rank and lower lighting sensitivity—and the future-feature probe localizes the advantage to changed patches at longer horizons. They also disclose limitations: the intervention evidence is qualitative, there are no uncertainty intervals, and the layer-selection probe is only 20 videos.\n\nThe main gap is mechanistic. The claim is that the reusable asset is the generator's internal computation, but the only evidence is Table 4's comparison against future pixels and action-only. Neither is a controlled substitute: pixels demand dense reconstruction, and action-only changes the gradient protocol. There is no baseline with a different future-predictive target (e.g., features from a V-JEPA or DINO-based future encoder) under the same architecture, the same stop-gradient task head, and the same compute. The paper also never reports how well the encoder actually predicts the multi-level generator target (G2R loss or cosine similarity on held-out clips). If the target is only weakly predictable, the gains could come from the strong base features plus the auxiliary objective rather than from internalizing generator computation. This does not destroy the paper—the efficiency result stands regardless—but it does mean the strongest conceptual claim is under-supported.\n\nMinor points: the layer set L={7,15,23,27} is post hoc on a 20-video probe, and the paper admits no universally best layer; the RoboTwin margin over Fast-WAM (0.19 points) is too small to support ranking without confidence intervals; and no code or data is released.\n\nWho this is for: anyone working on world models for control, VLA efficiency, or predictive representations. It deserves a serious referee. I would push for a controlled alternative-target baseline, reporting G2R prediction accuracy, and uncertainty estimates before accepting.","headline":"Solid empirical paper; the core efficiency claim holds, but the mechanism claim is under-tested because no controlled alternative future-predictive target is compared.","tokens_in":23641,"tokens_out":2033,"would_cite":true,"duration_ms":28122,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A world model's future-generating computation can be folded into a representation read from the present alone, making action prediction up to ten times faster without sacrificing task success.","keywords":["world models","predictive representation learning","robotic manipulation","visuomotor control","generative model distillation","action latency","flow matching","future prediction"],"falsifier":"Train the identical model with the multi-level generator-state target replaced by random, noise-matched vectors of the same shape while keeping all other losses and the downstream protocol fixed. If average success stays near 97.8%, the specific generator computation is not load-bearing; if it falls toward the action-only level (~94.9%), the generator-state target is what carries the claim.","tokens_in":22762,"feed_emoji":"🤖","tokens_out":8376,"duration_ms":111486,"temperature":0.7,"pith_summary":"The paper's central claim is that the most reusable asset of a generative world model is not the future video it renders but the internal computation that constructs that future, and that this computation can be transferred into a representation predicted from the current observation and instruction alone. Enfold does this by training a current-only encoder to predict the multi-level hidden states of a teacher-forced generator as it transforms a corrupted observed future into a coherent trajectory, then reading actions from a detached copy of the resulting representation. At deployment the generator is not executed for action prediction; the paper reports that this preserves or improves control accuracy on standard simulated manipulation benchmarks while cutting action latency by 3.7x, and by 10.1x with operator-level acceleration. If correct, the result changes what a world model is for: a training-time source of structured predictive supervision rather than a mandatory component of the control loop.","feed_headline":"World model folded into present: 10x faster actions","feed_subtitle":"Action heads read the encoded future-structure directly, so the generator runs only when imagination is requested.","key_machinery":"The load-bearing object is the multi-level generator-state target: internal features exposed at selected depths and corruption levels as a teacher-forced generator converts a corrupted latent of the real future into a coherent trajectory. These states are concatenated and layer-normalized, and a timestep-conditioned prediction head maps the current-only representation onto them (G2R). The representation is detached before conditioning future generation (R2G) and before task readouts, so task gradients cannot reshape the encoder. The mechanism makes the generator a training-time supervisor and optional decoder rather than an action-time component.","core_discovery":"Enfold claims that future-conditioned generative computation can be internalized in a current-only representation. During training, a video generator processes the observed future under corruption and exposes states at several depths; a timestep-conditioned head predicts these states from the current context and instruction, and a detached copy of the representation also conditions future generation and task readouts. At deployment, action prediction runs only the encoder and action head. The paper reports 97.8% and 91.77% average success on its two simulated suites, a 3.7x latency reduction (10.1x with operator-level acceleration) against the strongest world-action baseline considered, impr","pith_inferences":["A direct extension the paper does not pursue is learned selection of which generator depths supervise the encoder; its own probe shows the most informative layer shifts with corruption level, so a per-input or per-task routing could strengthen the G2R target.","The paper's projection interpretation suggests the G2R objective is an amortized conditional expectation; if so, the same scheme could be applied to any predictive teacher, not just a video generator, by exposing intermediate states of other future-constructing computations.","The token-geometry results point toward a testable hypothesis: the representation learns interaction relations (gripper-object) rather than object categories; one could probe this with a decodability experiment on relational predicates.","The 10.1x latency number includes operator-level acceleration, so the architectural contribution to speed is the 3.7x; separating the two matters when porting the method to a different runtime."],"forward_implications":["At control time, action prediction needs only the encoder and an action head, so world-model reasoning costs one forward pass instead of a full generative rollout.","The same representation is a functional input to future generation, so a robot can imagine a rollout only when needed while keeping the control loop cheap.","Supervision from multi-level generator states beats future-pixel and action-only supervision, and multi-level concatenation adds the largest gains on goal-directed and multi-stage tasks.","The representation suppresses nuisance variation from the generator (lighting sensitivity roughly 8-10x lower than raw generator features) while retaining more feature diversity, consistent with predictive filtering rather than feature collapse.","When the current scene is perturbed, both the imagined continuation and the executed actions redirect, indicating the policy is not replaying a fixed trajectory."],"fun_headline_variants":["Enfold: 10x faster actions, world model not run at inference","Fold future imagination into present, get 10x faster control","Enfold: Predict actions 10x faster by encoding the future","10x latency cut: Enfold distills world model into encoder","Enfold: 3.7x faster than Fast-WAM, 10x with acceleration"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The central premise is that the fixed set of generator depths chosen as supervision targets is both predictable from the current context and sufficient for control; the paper's own small probe shows no universally best layer, so the multi-level target could be mis-specified for tasks or corruption levels outside the probe.","fun_headline_variants_meta":{"raw":{"variants":["Enfold: 10x faster actions, world model not run at inference","Fold future imagination into present, get 10x faster control","Enfold: Predict actions 10x faster by encoding the future","10x latency cut: Enfold distills world model into encoder","Enfold: 3.7x faster than Fast-WAM, 10x with acceleration"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00027,"raw_usage":{"total_tokens":1490,"prompt_tokens":797,"completion_tokens":693,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":541,"completion_tokens_details":{"reasoning_tokens":595}},"tokens_in":541,"tokens_out":693,"duration_ms":9617,"temperature":1.0,"reasoning_tokens":595,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T03:15:42.384902+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Train the identical model with the multi-level generator-state target replaced by random, noise-matched vectors of the same shape while keeping all other losses and the downstream protocol fixed. If average success stays near 97.8%, the specific generator computation is not load-bearing; if it falls toward the action-only level (~94.9%), the generator-state target is what carries the claim.","supporting_citations":[],"review_version":2}