{"id":"19af52cf-a5f5-4423-8b60-2650d1559789","arxiv_id":"2607.06988","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"A meta-trained test-time memory lets frozen world-action models absorb unlabeled human videos and outperform in-context video conditioning on real multi-embodiment manipulation.","lead":"WAM-TTT steers a frozen robot world-action model by absorbing unlabeled human videos into a small adaptive memory at test time, without robot actions or fine-tuning. It offers a practical way to specify new robot behaviors from raw human play while keeping the foundation model intact.","discovery_kind":"new_method","skeptic_critique":{"model":"grok-4.5","headline":"The headline OOD gain may partly reflect human videos recorded in the target New scenes, so the frozen-WAM + human-only TTT claim is less pure than the abstract suggests.","rationale":"The reader correctly flags phase alignment and distribution drift as the soft assumption behind control-useful fast weights, and CONDITIONAL is the right shape given missing error bars, one task regression (Stamp Paper), and no code/data. I agree those limits matter, but the single most load-bearing concern for the stated strongest claim is slightly different: New results appear to use human videos collected in the evaluation households (Appendix B, Figures B.1–B.2), so the 46.2% figure is not a clean demonstration that raw human play from elsewhere steers a frozen WAM into unseen homes. That does not invalidate the method—meta-training + KVM + video-side TTT still beats ICL and co-training under matched human context, and E.2/E.3 give qualitative support for broader robustness—but it means the abstract/strongest-claim wording overstates the purity of the test-time human-only, cross-scene story. A cross-scene human-video ablation would settle whether the memory remains control-useful without scene-matched human footage. Until then, keep CONDITIONAL: accept-shaped if authors clarify the human-video collection protocol, add the cross-scene check or rephrase claims, and supply uncertainty/code as the reader asked. No change to REJECT; the empirical core is still coherent.","tokens_in":24836,"tokens_out":825,"duration_ms":7221,"concrete_test":"Re-evaluate the 9-task New protocol with a held-out human-video source: for each New household trial, adapt TTT only from human demos of the same skill recorded in a different scene (cubicle or another home), never the evaluation room. Report the same 25-trial progress table for WAM-TTT vs WAM-ICL and LDA. If average New progress falls near ICL/LDA and the 46.2 vs 7.1 gap collapses, the headline claim needs rephrasing to in-scene human memory rather than cross-scene human play.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"The strongest claim is that unlabeled human videos alone adapt a frozen WAM's TTT memory so New-household progress jumps from 7.1% (WAM-ICL) to 46.2% (WAM-TTT) without robot actions or task fine-tuning. That comparison is real, but the New protocol is not a pure test of transfer from out-of-scene human play. Appendix B states human demos for meta-training and for test-time TTT are recorded with a GoPro \"directly in the actual household environments that we later evaluate as the New setting,\" while robot data is cubicle-only. Section 3.3 then adapts fast weights on those in-scene human videos via L_vg + λ L_KVM. So the large New gap vs ICL partly measures whether in-scene human Key/Value memory helps more than long-context tokens under the same visual domain, not whether watching human play in a different place steers a frozen WAM into a new home. The paper's own E.2 lab-scene rollouts without in-scene human data are only qualitative; Table 1/C.1 New numbers all use in-scene human videos. The reader's phase-alignment concern is real (Limitations), but the more load-bearing issue for the central claim is this scene-matched human-video confound: control-usefulness of the meta-trained Q/K/V interface is demonstrated under human videos that already share New lighting/clutter/objects, which softens the \"steering by watching human play at test time\" framing.","agreement_with_reader":"partial"},"referee_report":{"model":"grok-4.5","summary":"WAM-TTT proposes a test-time training method to steer a frozen world-action model (built on LDA) using unlabeled human videos. A meta-training stage on phase-aligned paired human–robot data attaches video-side TTT residual branches and trains slow projections plus a key–value memory reconstruction loss so that human Keys/Values become a control-useful fast-weight memory; at deployment only those fast weights are updated by human-side video prediction and L_KVM while the WAM and action expert stay frozen. Real-robot evaluation on three embodiments and nine manipulation tasks reports large average gains in unseen household (New) settings over in-context human-video conditioning (WAM-ICL: 7.1% → 46.2%), the frozen backbone, co-training, and reimplemented EGOSCALE/π0.5 baselines, with ablations isolating meta-training, L_KVM, and TTT, plus data-ratio and pseudo-action studies.","tokens_in":25326,"tokens_out":1492,"duration_ms":19315,"significance":"If the result holds under a cleanly stated protocol, the paper offers a practical interface for RFM steering: absorb raw human play into a lightweight residual memory without robot actions, retargeting, or full fine-tuning, while keeping the foundation model frozen. The real-robot suite (3 embodiments, 9 tasks, Orig./New splits), the direct WAM-ICL control with the same human videos, and the informative ablations (especially pseudo-action harm and data-ratio iso-budget) are genuine strengths. The linear-attention witness in Appendix A is used only as motivation for L_KVM, not as a circular proof of performance. The work is a solid systems/methods contribution for world-action models and human-video transfer, contingent on clarifying how much of the New gain depends on scene-matched human videos.","major_comments":[{"comment":"Appendix B and Figures B.1–B.2 state that paired human demonstrations (and the human videos used for test-time TTT) are recorded with a GoPro “directly in the actual household environments that we later evaluate as the New setting,” while robot data is cubicle-only. Section 3.3 then adapts fast weights on those in-scene videos via L_vg + λ L_KVM. Table 1 / Table C.1 New numbers therefore compare TTT vs ICL under human videos that already share New lighting, clutter, and object instances—not pure transfer from out-of-scene human play. Appendix E.2’s “no in-scene human data” lab results are only qualitative. This is load-bearing for the abstract/intro framing of steering into new homes by watching human play. Please either (i) report quantitative New progress with human videos recorded outside the evaluation scene (or with cubicle-only human videos), or (ii) reframe claims and contribution","section":"§3.3, §4.1–4.2, App. B, App. E.2, Table 1/C.1"},{"comment":"Table 1 New average (46.2%) is driven by large wins on several tasks, but Stamp Paper is a clear failure (WAM-TTT 8.3 vs LDA 33.3). The text attributes this to tight stamp geometry and household perturbation, yet the paper still claims consistent outperformance “across diverse manipulation tasks.” Either provide a failure analysis (e.g., whether human videos lack the corrective cue, or L_KVM overwrites a useful prior) or qualify the consistency claim and discuss when human-video TTT can hurt relative to the frozen backbone.","section":"§4.2, Table 1, Table C.1"},{"comment":"Table 2 ablations (meta-training, memory recon., TTT, LoRA) use only two tasks and 10 trials per cell, while the main claim rests on nine tasks × 25 trials. Given that w/o Meta Training collapses on Swap Place (0.0) and WAM-LoRA is 0.0 there, the design isolation is important but under-powered. Extend the protocol ablation to at least the full New suite (or a larger fixed subset) with the same 25-trial protocol as Table 1, or report confidence intervals so the component contributions are not over-read from two tasks.","section":"§4.3, Table 2"}],"minor_comments":[{"comment":"No error bars or trial-level variance are reported for Table 1/C.1 despite 25 trials; adding mean±std or bootstrap intervals would make the +39.1 pt ICL gap easier to assess.","section":"§4.2, Table 1, Table C.1"},{"comment":"Eq. (3)–(5) and Table B.1: N=1 inner SGD step is aggressive; a short sensitivity note on N and η_test would help readers judge stability of the fast-weight update.","section":"§3.2–3.3, Table B.1"},{"comment":"Figure 1 / title use “W AM” / “WAM-TTT” spacing inconsistently; unify notation (WAM vs W AM) throughout.","section":"Title, Figure 1, Abstract"},{"comment":"Appendix D Stamp Paper rubric has a typo: “stamp s[uccessfully grasped.”","section":"App. D"},{"comment":"Related work on TTT and human-video transfer is thorough; a one-sentence contrast with MimicDroid [18] in the main text (beyond the citation list) would clarify the ICL baseline choice.","section":"§2, §4.2"},{"comment":"Table 3 generalization results are only on Deliver Drink; stating that scope in the caption would avoid over-generalizing “all perturbation types.”","section":"§4.4, Table 3"}],"recommendation":"major_revision","confidential_remarks":"The technical idea (video-side TTT memory + KVM meta-alignment) is interesting and the robot evidence is above average for this area. The main risk for the journal is overclaim relative to the scene-matched human-video protocol; if the authors reframe cleanly or add the out-of-scene quantitative control, this is a strong methods paper. I would not reject on novelty grounds alone—the ICL vs TTT comparison with the same human videos is a useful contribution even under the current protocol."},"author_rebuttal":null,"desk_editor":{"model":"grok-4.5","letter":"Punchline: this is a clean empirical methods paper that actually ships a usable idea—meta-train a video-side TTT branch with a key–value reconstruction loss so that, at deploy time, only unlabeled human RGB updates fast weights inside a frozen world-action model. The ICL comparison is the right one, and the gap is large (46.2% vs 7.1% New average). That is the contribution worth knowing.\n\nWhat is new is not TTT, WAMs, or human-video IL separately, but the calibrated residual memory: human Keys/Values written by video prediction + L_KVM, robot Queries reading them, backbone frozen. Real robots across three embodiments and nine tasks, plus ablations that isolate meta-training, KVM, and TTT, make this more than a recipe dump. The pseudo-action ablation is especially honest: adding MANO retargeting and FD hurts hard, which supports keeping human video action-free. Data-ratio results also land—paired human data substitutes for robot demos at fixed budget without magic.\n\nSoft spots, in proportion. The stress-test is right on the main framing: Appendix B says human demos for meta-training and test-time TTT are recorded in the actual New household scenes, while robot data is cubicle-only. So Table 1’s New win is “in-scene human Key/Value memory beats long-context tokens under the same visual domain,” not pure transfer from watching play in a different place. E.2 (no in-scene human video) is only qualitative. That does not kill the method—ICL still collapses under the same protocol—but it softens the abstract’s “steering by watching human play” claim and should be stated up front. Phase-alignment dependence is already in Limitations; treat it as a real constraint. No error bars, Stamp Paper regresses vs LDA, free hyperparameters (λ, N, η, head dims), no code/data release. Math in App. A is a linear-attention witness for L_KVM, not a circular proof of the robot results.\n\nWho this is for: people building steerable RFMs / WAMs who care about deployment-time human interface without retargeting or full fine-tuning. Worth a serious referee. I would engage, cite if I work on this stack, and bring it to reading group with the scene-matched human-video caveat on the board.","headline":"Solid systems paper: TTT memory for frozen WAMs from action-free human video is real and useful, but the big New-household numbers use in-scene human demos, so the pure “watch play elsewhere, steer here” story is softer than the abstract sells.","tokens_in":25968,"tokens_out":623,"would_cite":true,"duration_ms":9378,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.5","headline":"Unlabeled human play videos can steer a frozen robot foundation model at test time by adapting only a lightweight video-side memory.","keywords":["world action models","test-time training","human videos","robot foundation models","manipulation","fast-weight memory","key-value reconstruction"],"falsifier":"On the same nine real-robot tasks and unseen household setting, replace the test-time fast-weight update with pure in-context conditioning on the identical human videos (or remove meta-training / the key–value loss) and check whether average progress collapses back toward the reported 7.1% baseline instead of remaining near 46%.","tokens_in":25720,"feed_emoji":"🤖","tokens_out":589,"duration_ms":4817,"temperature":0.7,"pith_summary":"Robot foundation models are hard to steer toward new task variants or user-preferred behaviors without more robot demos, full fine-tuning, or long context. This paper claims that raw human videos need not be treated as trajectories to imitate. Instead they can be absorbed into a small adaptive memory inside a frozen world-action model through self-supervised video prediction. A prior meta-training stage on paired human–robot data aligns that memory so human visual cues become useful for robot control via a key–value reconstruction objective. At deployment only unlabeled human videos update the memory; the backbone stays frozen. The result is efficient, reusable steering that preserves the foundation model’s generalization and, on real multi-embodiment manipulation, substantially outperforms feeding the same videos as in-context conditioning.","feed_headline":"Human play videos steer frozen robot models at test time","feed_subtitle":"A lightweight memory absorbs unlabeled demos; the foundation model stays frozen and still generalizes.","key_machinery":"WAM-TTT: residual TTT (test-time training) layers on the video expert of a frozen world-action model. Fast weights absorb human videos via video prediction plus key–value memory reconstruction; robot Queries read the adapted memory as a residual that steers action generation through shared visual-action dynamics.","core_discovery":"A world-action model can be steered at test time from action-free human videos alone by updating only a lightweight fast-weight memory on the video expert, provided that memory was first meta-trained with paired human–robot data and a key–value reconstruction loss so that human Keys/Values become control-useful residuals for robot Queries.","pith_inferences":[],"forward_implications":[],"fun_headline_variants":["Test-time memory steers frozen robots from raw human play videos","Lightweight adaptive memory absorbs human demos for robot control","Human play videos update fast weights to steer frozen WAMs","Meta-trained memory turns unlabeled human videos into robot cues","Steer world-action models at test time by watching human play"],"cache_read_input_tokens":16512,"weakest_assumption_plain":"That phase-aligned paired human–robot meta-training produces a memory interface that stays useful for control when the only later updates come from unlabeled human videos on new tasks and scenes.","fun_headline_variants_meta":{"raw":{"variants":["Test-time memory steers frozen robots from raw human play videos","Lightweight adaptive memory absorbs human demos for robot control","Human play videos update fast weights to steer frozen WAMs","Meta-trained memory turns unlabeled human videos into robot cues","Steer world-action models at test time by watching human play"]},"model":"grok-4.5","effort":"low","cost_usd":0.003686,"raw_usage":{"total_tokens":1165,"prompt_tokens":730,"num_sources_used":0,"completion_tokens":85,"cost_in_usd_ticks":36860000,"prompt_tokens_details":{"text_tokens":730,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":350,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":730,"tokens_out":85,"duration_ms":3432,"temperature":1.0,"reasoning_tokens":350,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-13T06:47:13.977193+00:00","model_set":{"reader":"grok-4.5"},"falsifier":"On the same nine real-robot tasks and unseen household setting, replace the test-time fast-weight update with pure in-context conditioning on the identical human videos (or remove meta-training / the key–value loss) and check whether average progress collapses back toward the reported 7.1% baseline instead of remaining near 46%.","supporting_citations":[],"review_version":2}