Pith. sign in

REVIEW 3 major objections 6 minor 1 cited by

WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time

T0 review · 3 major / 6 minor · reviewed 2026-07-13 · grok-4.5

Pith's one-line read Unlabeled human play videos can steer a frozen robot foundation model at test time by adapting only a lightweight video-side memory.

desk verdict Solid systems paper: TTT memory for frozen WAMs from action-free human video is real and useful, but the big New-household numbers use in-scene human demos, so the pure “watch play elsewhere, steer here” story is softer than the abstract sells. read the letter →

arxiv 2607.06988 v2 pith:6VH44RPU submitted 2026-07-08 cs.RO cs.AI

classification cs.ROcs.AI
keywords worldactionmodelstest-timetraininghumanvideosrobotfoundationmanipulationfast-weightmemorykey-valuereconstruction
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

Robot foundation models are hard to steer toward new task variants or user-preferred behaviors without more robot demos, full fine-tuning, or long context. This paper claims that raw human videos need not be treated as trajectories to imitate. Instead they can be absorbed into a small adaptive memory inside a frozen world-action model through self-supervised video prediction. A prior meta-training stage on paired human–robot data aligns that memory so human visual cues become useful for robot control via a key–value reconstruction objective. At deployment only unlabeled human videos update the memory; the backbone stays frozen. The result is efficient, reusable steering that preserves the foundation model’s generalization and, on real multi-embodiment manipulation, substantially outperforms feeding the same videos as in-context conditioning.

What carries the argument

WAM-TTT: residual TTT (test-time training) layers on the video expert of a frozen world-action model. Fast weights absorb human videos via video prediction plus key–value memory reconstruction; robot Queries read the adapted memory as a residual that steers action generation through shared visual-action dynamics.

What would settle it

On the same nine real-robot tasks and unseen household setting, replace the test-time fast-weight update with pure in-context conditioning on the identical human videos (or remove meta-training / the key–value loss) and check whether average progress collapses back toward the reported 7.1% baseline instead of remaining near 46%.

Watch

Extended reading notes

Core claim

A world-action model can be steered at test time from action-free human videos alone by updating only a lightweight fast-weight memory on the video expert, provided that memory was first meta-trained with paired human–robot data and a key–value reconstruction loss so that human Keys/Values become control-useful residuals for robot Queries.

Load-bearing premise

That phase-aligned paired human–robot meta-training produces a memory interface that stays useful for control when the only later updates come from unlabeled human videos on new tasks and scenes.

Share X Bluesky LinkedIn Reddit HN

Signed reviews

No signed human review yet.

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

3 major / 6 minor

Summary. WAM-TTT proposes a test-time training method to steer a frozen world-action model (built on LDA) using unlabeled human videos. A meta-training stage on phase-aligned paired human–robot data attaches video-side TTT residual branches and trains slow projections plus a key–value memory reconstruction loss so that human Keys/Values become a control-useful fast-weight memory; at deployment only those fast weights are updated by human-side video prediction and L_KVM while the WAM and action expert stay frozen. Real-robot evaluation on three embodiments and nine manipulation tasks reports large average gains in unseen household (New) settings over in-context human-video conditioning (WAM-ICL: 7.1% → 46.2%), the frozen backbone, co-training, and reimplemented EGOSCALE/π0.5 baselines, with ablations isolating meta-training, L_KVM, and TTT, plus data-ratio and pseudo-action studies.

Significance. If the result holds under a cleanly stated protocol, the paper offers a practical interface for RFM steering: absorb raw human play into a lightweight residual memory without robot actions, retargeting, or full fine-tuning, while keeping the foundation model frozen. The real-robot suite (3 embodiments, 9 tasks, Orig./New splits), the direct WAM-ICL control with the same human videos, and the informative ablations (especially pseudo-action harm and data-ratio iso-budget) are genuine strengths. The linear-attention witness in Appendix A is used only as motivation for L_KVM, not as a circular proof of performance. The work is a solid systems/methods contribution for world-action models and human-video transfer, contingent on clarifying how much of the New gain depends on scene-matched human videos.

major comments (3)
  1. [§3.3, §4.1–4.2, App. B, App. E.2, Table 1/C.1] Appendix B and Figures B.1–B.2 state that paired human demonstrations (and the human videos used for test-time TTT) are recorded with a GoPro “directly in the actual household environments that we later evaluate as the New setting,” while robot data is cubicle-only. Section 3.3 then adapts fast weights on those in-scene videos via L_vg + λ L_KVM. Table 1 / Table C.1 New numbers therefore compare TTT vs ICL under human videos that already share New lighting, clutter, and object instances—not pure transfer from out-of-scene human play. Appendix E.2’s “no in-scene human data” lab results are only qualitative. This is load-bearing for the abstract/intro framing of steering into new homes by watching human play. Please either (i) report quantitative New progress with human videos recorded outside the evaluation scene (or with cubicle-only human videos), or (ii) reframe claims and contribution
  2. [§4.2, Table 1, Table C.1] Table 1 New average (46.2%) is driven by large wins on several tasks, but Stamp Paper is a clear failure (WAM-TTT 8.3 vs LDA 33.3). The text attributes this to tight stamp geometry and household perturbation, yet the paper still claims consistent outperformance “across diverse manipulation tasks.” Either provide a failure analysis (e.g., whether human videos lack the corrective cue, or L_KVM overwrites a useful prior) or qualify the consistency claim and discuss when human-video TTT can hurt relative to the frozen backbone.
  3. [§4.3, Table 2] Table 2 ablations (meta-training, memory recon., TTT, LoRA) use only two tasks and 10 trials per cell, while the main claim rests on nine tasks × 25 trials. Given that w/o Meta Training collapses on Swap Place (0.0) and WAM-LoRA is 0.0 there, the design isolation is important but under-powered. Extend the protocol ablation to at least the full New suite (or a larger fixed subset) with the same 25-trial protocol as Table 1, or report confidence intervals so the component contributions are not over-read from two tasks.
minor comments (6)
  1. [§4.2, Table 1, Table C.1] No error bars or trial-level variance are reported for Table 1/C.1 despite 25 trials; adding mean±std or bootstrap intervals would make the +39.1 pt ICL gap easier to assess.
  2. [§3.2–3.3, Table B.1] Eq. (3)–(5) and Table B.1: N=1 inner SGD step is aggressive; a short sensitivity note on N and η_test would help readers judge stability of the fast-weight update.
  3. [Title, Figure 1, Abstract] Figure 1 / title use “W AM” / “WAM-TTT” spacing inconsistently; unify notation (WAM vs W AM) throughout.
  4. [App. D] Appendix D Stamp Paper rubric has a typo: “stamp s[uccessfully grasped.”
  5. [§2, §4.2] Related work on TTT and human-video transfer is thorough; a one-sentence contrast with MimicDroid [18] in the main text (beyond the citation list) would clarify the ICL baseline choice.
  6. [§4.4, Table 3] Table 3 generalization results are only on Deliver Drink; stating that scope in the caption would avoid over-generalizing “all perturbation types.”

Circularity Check

0 steps flagged · score 0.0 of 10

No circular derivation: empirical TTT method whose control claims rest on held-out robot progress, not on restating fitted inputs or self-citation uniqueness.

full rationale

WAM-TTT is a methods paper: meta-train a TTT branch on paired human–robot data (outer L_robot_WAM + inner L_vg + λ L_KVM), then at test time update only fast weights from unlabeled human videos while freezing the WAM. Performance claims (e.g., 46.2% vs 7.1% New progress vs WAM-ICL) are measured on real-robot trials with external baselines (π0.5, EGOSCALE, LDA), not obtained by renaming a fit as a prediction. Appendix A’s linear-attention “witness” is a closed-form motivation for L_KVM in the linear special case; it does not force the empirical robot results by construction. Citations to LDA [23] and Spatial-TTT [52] supply the backbone and TTT layer form—normal scaffolding, not a load-bearing uniqueness theorem that forbids alternatives. Phase-alignment and in-scene human-video protocol issues (if any) are experimental confounds, not circular reductions of equations to their inputs. No self-definitional loop, fitted-input-as-prediction, or ansatz-smuggled uniqueness chain is present. Score 0 with empty steps is the honest finding.

Assumptions & free parameters 5 free parameters · 5 assumptions · 2 invented entities

The central claim rests on standard deep-learning and robotics assumptions plus several paper-specific design choices: that a residual fast-weight memory on the video expert can carry human skill into robot actions; that phase-aligned paired data is enough to learn that interface; and that a handful of free hyperparameters (inner steps, LRs, λ, TTT width) generalize across tasks. No new physical entities are postulated; the invented pieces are architectural/objective constructs whose only evidence is the paper’s robot trials.

free parameters (5)
  • memory reconstruction weight λ = 4e-2
    Balances L_KVM against video prediction in the inner loop; set to 4e-2 in Table B.1 and used at both meta-train and test time.
  • inner SGD steps N = 1
    Number of fast-weight updates per example/deployment; fixed to 1 for both stages.
  • inner learning rates η_meta / η_test = 0.1 / 0.01
    Hand-chosen rates for fast-weight adaptation (0.1 meta-training, 0.01 test time).
  • TTT head dim d and fast-weight hidden width = 48 / 128
    Capacity of the residual memory network; set to 48 / 128 without a full capacity sweep in the main text.
  • meta-training data mix (robot, human) per task = (100, 100)
    Default (100,100) chosen after a limited data-ratio ablation; performance depends on this budget and mix.
assumptions (5)
  • domain assumption A residual fast-weight update on video tokens can steer action generation through joint video-action attention without updating the action expert.
    Core architectural premise of §3.1 and Eq. 1–2; justified by WAM coupling but not independently proven.
  • ad hoc to paper Nearest-phase synchronization of human and robot episodes yields a valid human–robot alignment signal for meta-training.
    Section 3.2 and Limitations: mis-aligned phase distributions degrade the inner adaptation signal.
  • ad hoc to paper Minimizing key–value reconstruction on human Keys/Values produces a memory that robot Queries can usefully read (linear-attention witness).
    Appendix A derives the linear case; the deployed model is a nonlinear MLP, so the witness is motivational rather than exact.
  • domain assumption Self-supervised video prediction on unlabeled human videos is a sufficient test-time objective for control-useful memory when the Q/K/V interface was meta-trained.
    Section 3.3; supported by ablations but assumed to transfer beyond the meta-training task distribution.
  • domain assumption Standard diffusion / flow-matching multitask losses from the LDA backbone are valid outer objectives for joint WAM + TTT training.
    Inherited from LDA [23] without modification (Eq. 6).
invented entities (2)
  • Video-side TTT residual branches as human skill memory inside a frozen WAM
    purpose: Store deployment-time human video information in fast weights without long context or full fine-tuning.
    Architectural construct introduced in §3.1; evidence is only the paper’s robot experiments.
  • Key–value memory reconstruction loss L_KVM for human–robot meta-alignment
    purpose: Force fast weights to reconstruct human Values from human Keys so robot Queries can read human-derived residuals.
    Paper-specific objective (Eq. 3); ablations show removing it hurts, but no external independent validation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time." pith.science (2026). https://pith.science/paper/6VH44RPU

@misc{pith2026260706988,
  author       = {Pith},
  title        = {Pith review of: WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/6VH44RPU}},
  note         = {Machine review of arXiv:2607.06988}
}
read the original abstract

Steering robot foundation models (RFMs) toward new task variants or user-preferred behaviors remains challenging, often requiring additional robot demonstrations, task-specific fine-tuning, or long-context conditioning. We present WAM-TTT, a test-time training framework for steering world action models from raw human videos. Rather than treating human videos as trajectories to imitate, WAM-TTT absorbs them into a lightweight adaptive memory inside a frozen WAM through self-supervised video prediction. To make this memory useful for control, we introduce a meta-training stage that aligns human demonstrations with robot behaviors using paired human-robot data and a key--value memory reconstruction objective. At test time, only unlabeled human videos are required to adapt the memory, while the pretrained WAM remains frozen. This enables efficient and reusable steering without robot actions, human-side annotations, or task-specific fine-tuning, while preserving the generalization ability of the foundation model. Extensive experiments show that WAM-TTT consistently outperforms in-context human-video conditioning baselines across diverse manipulation tasks and generalization settings.

Figures

Figures reproduced from arXiv: 2607.06988 by the authors.

Figure 1
Figure 1. Overview of WAM-TTT. Given unlabeled human demonstrations from diverse environ￾ments, WAM-TTT steers a pretrained World Action Model (WAM) without retargeting, robot actions, or human-side annotations. During deployment, human videos are absorbed into lightweight TTT fast weights through self-supervised video prediction, while the pretrained action model remains frozen. The adapted memory then guides robot execution… view at source ↗
Figure 2
Figure 2. Pipeline of WAM-TTT. We first meta-train a fast-weight memory using paired human-robot demonstrations, encouraging human visual cues to align with robot behaviors through a key–value memory reconstruction objective. At test time, the memory is adapted from unlabeled human videos via video prediction, while the pretrained WAM remains frozen. The adapted memory then steers robot execution through the WAM’s shared visu… view at source ↗
Figure 3
Figure 3. Experimental setup [PITH_FULL_IMAGE:figures/full_fig_p006_3.png] view at source ↗
Figures from the paper (2 more)
Figure 4
Figure 4. Figure 4: Qualitative rollouts. For each unseen task we show a robot rollout filmstrip (right) and the paired human demonstration used as deployment-time Key/Value (left). Dataset and Metric. We collect a meta-training dataset consisting of 2,286 paired human and robot episodes,…
Figure 5
Figure 5. Figure 5: Generalization Setup. including lighting, object position, and embodiment-related appearance shifts [PITH_FULL_IMAGE:figures/full_fig_p008_5.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models

    cs.RO 2026-08 conditional novelty 5.0 of 10

    A vision-language-action policy conditioned on one automatically structured demonstration (sub-goals plus verbalized 3D/2D motion) achieves top scores on LIBERO, LIBERO-Plus, and VLA-Arena without fine-tuning.

Reference graph

Works this paper leans on

61 extracted references · 30 linked inside Pith · cited by 1 Pith paper

  1. [1]

    Yu et al

    T. Yu et al. One-shot imitation from observing humans via domain-adaptive meta-learning. In RSS, 2018

  2. [2]

    S. Bahl, A. Gupta, and D. Pathak. Human-to-robot imitation in the wild. InRSS, 2022

  3. [3]

    M. Xu, Z. Xu, C. Chi, M. Veloso, and S. Song. Xskill: Cross embodiment skill discovery. In CoRL, 2023

  4. [4]

    Bharadhwaj, A

    H. Bharadhwaj, A. Gupta, V . Kumar, and S. Tulsiani. Towards generalizable zero-shot manip- ulation via translating human interaction plans. InICRA, 2024

  5. [5]

    Hansen et al

    N. Hansen et al. Self-supervised policy adaptation during deployment. InICLR, 2021

  6. [6]

    M. Xu, Z. Xu, C. C. Pan, X. Zhu, C. Tomei, Y . Shen, Z. Wu, S.-R. Chen, J. B. Tenenbaum, T. Lozano-Perez, and S. Song. Flow as the cross-domain manipulation interface. InCoRL, 2024

  7. [7]

    Kareer, D

    S. Kareer, D. Patel, R. Punamiya, P. Mathur, S. Cheng, C. Wang, J. Hoffman, and D. Xu. Egomimic: Scaling imitation learning via egocentric video.arXiv preprint arXiv:2410.24221, 2024

  8. [8]

    Hoque, P

    R. Hoque, P. Huang, D. J. Yoon, M. Sivapurapu, and J. Zhang. Egodex: Learning dexterous manipulation from large-scale egocentric video.arXiv preprint arXiv:2505.11709, 2025

Show all 61 references
  1. [9]

    Grauman et al

    K. Grauman et al. Ego-exo4d: Understanding skilled human activity from first- and third- person perspectives. InCVPR, 2024

  2. [10]

    Zheng et al

    R. Zheng et al. Egoscale: Scaling dexterous manipulation with diverse egocentric human data. arXiv preprint arXiv:2602.16710, 2026

  3. [11]

    Chen et al

    H. Chen et al. Vidbot: Learning generalizable 3d actions from in-the-wild 2d human videos for zero-shot robotic manipulation. InCVPR, 2025

  4. [12]

    Kim et al

    H. Kim et al. Uniskill: Imitating human videos via cross-embodiment skill representations. In CoRL, 2025

  5. [13]

    Z. Chen, S. Chen, E. Arlaud, I. Laptev, and C. Schmid. Vividex: Learning vision-based dex- terous manipulation from human videos. InICRA, 2025

  6. [14]

    C. Wang, L. Fan, J. Sun, R. Zhang, L. Fei-Fei, D. Xu, Y . Zhu, and A. Anandkumar. Mimicplay: Long-horizon imitation learning by watching human play. InCoRL, 2023. arXiv:2302.12422

  7. [15]

    Bharadhwaj, A

    H. Bharadhwaj, A. Gupta, S. Tulsiani, and V . Kumar. Zero-shot robot manipulation from passive human videos.arXiv preprint arXiv:2302.02011, 2023

  8. [16]

    V . Jain, M. Attarian, N. J. Joshi, A. Wahid, D. Driess, Q. Vuong, P. R. Sanketi, P. Ser- manet, S. Welker, C. Chan, I. Gilitschenski, Y . Bisk, and D. Dwibedi. Vid2robot: End- to-end video-conditioned policy learning with cross-attention transformers.arXiv preprint arXiv:2403...

  9. [17]

    Bharadhwaj, D

    H. Bharadhwaj, D. Dwibedi, A. Gupta, S. Tulsiani, C. Doersch, T. Xiao, D. Shah, F. Xia, D. Sadigh, and S. Kirmani. Gen2act: Human video generation in novel scenarios enables generalizable robot manipulation.arXiv preprint arXiv:2409.16283, 2024

  10. [18]

    R. Shah, S. Liu, Q. Wang, Z. Jiang, S. Kumar, M. Seo, R. Mart ´ın-Mart´ın, and Y . Zhu. Mim- icdroid: In-context learning for humanoid robot manipulation from human play videos.arXiv preprint arXiv:2509.09769, 2025. 10

  11. [19]

    Chi, C.-K

    X. Chi, C.-K. Fan, H. Zhang, X. Qi, R. Zhang, A. Chen, C.-m. Chan, W. Xue, Q. Liu, S. Zhang, et al. Eva: An embodied world model for future video anticipation.arXiv preprint arXiv:2410.15461, 2024

  12. [20]

    X. Chi, P. Jia, C.-K. Fan, X. Ju, W. Mi, K. Zhang, Z. Qin, W. Tian, K. Ge, H. Li, et al. Wow: Towards a world omniscient world model through embodied interaction.arXiv preprint arXiv:2509.22642, 2025

  13. [21]

    Zhang, X

    J. Zhang, X. Chen, A.-J. Chen, C. Lv, D. mei Li, G. Zhou, H. Yin, H. Yuan, H. Li, J. Li, J. Zhang, J. Zhou, K. Gao, K. Yan, L. Jiang, N. Tang, P. Lin, Q. Peng, S.-S. Yin, T. Wu, T. Yan, X. Xu, Y . Shu, Y . Zhang, Y . Wang, Y . Wang, Y . Chen, Y . Xu, Y . Huang, Y . Chen, Z. Zh...

  14. [22]

    C. Zhu, R. Yu, S. Feng, B. Burchfiel, P. Shah, and A. Gupta. Unified world models: Cou- pling video and action diffusion for pretraining on large robotic datasets.arXiv preprint arXiv:2504.02792, 2025

  15. [23]

    J. Lyu, K. Liu, X. Zhang, H. Liao, Y . Feng, W. Zhu, T. Shen, J. Chen, J. Zhang, Y . Dong, W. Cui, S. Qi, S. Wang, Y . Zheng, M. Yan, X. Shi, H. Li, D. Zhao, M.-Y . Liu, Z. Zhang, L. Yi, Y . Wang, and H. Wang. LDA-1B: Scaling latent dynamics action model via universal embodied...

  16. [24]

    H. Bi, H. Tan, S. Xie, Z. Wang, S. Huang, H. Liu, R. Zhao, Y . Feng, C. Xiang, Y . Rong, et al. Motus: A unified latent action world model. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 35101–35113, 2026

  17. [25]

    M. Team, C. Xiang, F. Bao, H. Liu, H. Tan, H. Bi, J. Li, J. Liu, J. Pang, K. Jing, et al. Mo- tubrain: An advanced world action model for robot control.arXiv preprint arXiv:2604.27792, 2026

  18. [26]

    Zhang, W

    Y . Zhang, W. Zhang, Z. Qi, H. Zhang, H. Lin, J. Zhang, Y . Mu, X. Yang, W. Zeng, and X. Jin. Imagewam: Do world action models really need video generation, or just image editing?, 2026. URLhttps://arxiv.org/abs/2606.19531

  19. [27]

    H. Yu, H. Lin, J. Zhang, W. Zhang, C. Gu, H. Li, and P. Tan. Maskwam: Unifying mask prompting and prediction for world-action models.arXiv preprint arXiv:2606.13515, 2026

  20. [28]

    B. Peng, W. Zhang, L. Xu, Z. Qi, J. Zhang, H. Liu, W. Zeng, and X. Jin. Reworld: Multi- dimensional reward modeling for embodied world models.arXiv preprint arXiv:2601.12428, 2026

  21. [29]

    S. Ye, Y . Ge, K. Zheng, S. Gao, S. Yu, G. Kurian, S. Indupuru, Y . L. Tan, C. Zhu, J. Xi- ang, A. Malik, K. Lee, W. Liang, N. Ranawaka, J. Gu, Y . Xu, G. Wang, F. Hu, A. Narayan, J. Bjorck, J. Wang, G. Kim, D. Niu, R. Zheng, Y . Xie, J. Wu, Q. Wang, R. Julian, D. Xu, Y . Du, ...

  22. [30]

    A. Ye, B. Wang, C. Ni, G. Huang, G. Zhao, H. Li, H. Li, J. Li, J. Lv, J. Liu, M. Cao, P. Li, Q. Deng, W. Mei, X. Wang, X. Chen, X. Zhou, Y . Wang, Y . Chang, Y . Li, Y . Zhou, Y . Ye, Z. Liu, and Z. Zhu. Gigaworld-policy: An efficient action-centered world-action model.arXiv p...

  23. [31]

    L. Li, Q. Zhang, Y . Luo, S. Yang, R. Wang, F. Han, M. Yu, Z. Gao, N. Xue, X. Zhu, Y . Shen, and Y . Xu. Causal world modeling for robot control.arXiv preprint arXiv:2601.21998, 2026. 11

  24. [32]

    T. Yuan, Z. Dong, Y . Liu, and H. Zhao. Fast-wam: Do world action models need test-time future imagination?arXiv preprint arXiv:2603.16666, 2026. URLhttps://arxiv.org/ abs/2603.16666

  25. [33]

    J. Cai, L. Ling, S. Chu, Z. Liu, J. Kang, Z. Liang, W. Xu, Y . Mao, W. Zhang, X. Yang, R. Ying, R. Zheng, and Y . Mu. Aha-wam: Asynchronous horizon-adaptive world-action modeling with observation-guided context routing.arXiv preprint arXiv:2606.09811, 2026

  26. [34]

    Q. Feng, J. Yu, J. Liu, Y . Jia, Z. Wu, H. Chen, Z. Qian, S. Gu, P. Jia, S. Ma, and S. Zhang. Harmowam: Harmonizing generalizable and precise manipulation via adaptive world action models, 2026

  27. [35]

    H. Luo, W. Zhang, Y . Feng, S. Zheng, H. Xu, C. Xu, Z. Xi, Y . Fu, and Z. Lu. Being-h0. 7: A latent world-action model from egocentric videos.arXiv preprint arXiv:2605.00078, 2026

  28. [36]

    J. Guo, Q. Li, P. Li, Z. Chen, N. Sun, Y . Su, H. Wang, Y . Zhang, X. Li, and H. Liu. Unified 4d world action modeling from video priors with asynchronous denoising.arXiv preprint arXiv:2604.26694, 2026

  29. [37]

    J. Lyu, Z. Li, X. Shi, C. Xu, Y . Wang, and H. Wang. Dywa: Dynamics-adaptive world action model for generalizable non-prehensile manipulation.arXiv preprint arXiv:2503.16806, 2025

  30. [38]

    Agarwal, A

    N. Agarwal, A. Ali, J. Allen, M. Antolini, A. Aubame, A. Azzolini, J. Bai, M. Bala, Y . Bal- aji, J. Bapst, et al. Cosmos 3: Omnimodal world models for physical ai.arXiv preprint arXiv:2606.02800, 2026

  31. [39]

    T. Ma, J. Zheng, Z. Wang, C. Jiang, A. Cui, J. Liang, and S. Yang. Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control.arXiv preprint arXiv:2603.10448, 2026

  32. [40]

    Physical Intelligence, B. Ai, A. Amin, R. J. Aniceto, A. Balakrishna, G. Balke, K. Black, G. Bokinsky, S. Cao, T. Charbonnier, et al.π 0.7: A steerable generalist robotic foundation model with emergent capabilities.arXiv preprint arXiv:2604.15483, 2026

  33. [41]

    Liu et al

    Y . Liu et al. Oa-wam: Object-addressable world action model for robust robot manipulation. 2026

  34. [42]

    Y . Yang, S. Zeng, T. Lin, X. Chang, D. Qi, J. Xiao, H. Liu, R. Chen, Y . Chen, D. Huo, et al. Abot-m0: Vla foundation model for robotic manipulation with action manifold learning.arXiv preprint arXiv:2602.11236, 2026

  35. [43]

    R. Chen, Y . Yang, Z. Tang, D. Huo, T. Lin, H. Wu, H. Liu, Y . Chen, L. Zheng, B. Yuan, T. Li, M. Wang, D. Qi, B. Hu, W. Mei, Y . Xuan, H. Yang, Y . Zhu, M. Xu, Z. Ma, and X. Chang. Abot-m0.5: Unified mobility-and-manipulation world action model.arXiv preprint arXiv:2607.00678, 2026

  36. [44]

    M. J. Kim, Y . Gao, T.-Y . Lin, Y .-C. Lin, Y . Ge, G. Lam, P. Liang, S. Song, M.-Y . Liu, C. Finn, and J. Gu. Cosmos policy: Fine-tuning video models for visuomotor control and planning. In International Conference on Learning Representations (ICLR), 2026

  37. [45]

    Y . Hu, Y . Guo, P. Wang, X. Chen, Y .-J. Wang, J. Zhang, K. Sreenath, C. Lu, and J. Chen. Video prediction policy: A generalist robot policy with predictive visual representations.arXiv preprint arXiv:2412.14803, 2024

  38. [46]

    Y . Sun, X. Wang, Z. Liu, J. Miller, A. A. Efros, and M. Hardt. Test-time training with self- supervision for generalization under distribution shifts. InICML, 2020

  39. [47]

    D. Wang, E. Shelhamer, S. Liu, B. Olshausen, and T. Darrell. Tent: Fully test-time adaptation by entropy minimization. InICLR, 2021. 12

  40. [48]

    Y . Liu, P. Kothari, B. van Delft, B. Bellot-Gurlet, T. Mordan, and A. Alahi. Ttt++: When does self-supervised test-time training fail or thrive? InNeurIPS, 2021

  41. [49]

    Gandelsman, Y

    Y . Gandelsman, Y . Sun, X. Chen, and A. A. Efros. Test-time training with masked autoen- coders. InNeurIPS, 2022

  42. [50]

    Y . Sun, X. Li, K. Dalal, J. Xu, A. Vikram, G. Zhang, Y . Dubois, X. Chen, X. Wang, S. Koyejo, T. Hashimoto, and C. Guestrin. Learning to (learn at test time): Rnns with expressive hidden states. InICML, 2025

  43. [51]

    Behrouz, P

    A. Behrouz, P. Zhong, and V . Mirrokni. Titans: Learning to memorize at test time.arXiv preprint arXiv:2501.00663, 2025

  44. [52]

    F. Liu, D. Wu, J. Chi, Y . Cai, Y .-H. Hung, X. Yu, H. Li, H. Hu, Y . Rao, and Y . Duan. Spatial-ttt: Streaming visual-based spatial intelligence with test-time training.arXiv preprint arXiv:2603.12255, 2026

  45. [53]

    S. Yang, Y . Ze, and H. Xu. Movie: Visual model-based policy adaptation for view generaliza- tion. InNeurIPS, 2023

  46. [54]

    Z. Bai, C. Gao, and M. Z. Shou. Evolve-vla: Test-time training from environment feedback for vision-language-action models.arXiv preprint arXiv:2512.14666, 2025

  47. [55]

    C. Liu, Y . Liu, T. Wang, Q. Zhuang, J. C. Liang, W. Yang, R. Xu, Q. Wang, D. Liu, and C. Han. On-the-fly vla adaptation via test-time reinforcement learning.arXiv preprint arXiv:2601.06748, 2026

  48. [56]

    Black, N

    K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Haus- man, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. Vuong, H. Wal...

  49. [57]

    J. Mu, S. Yang, Y . Bao, H. Bae, T. Wei, L. Xu, B. Li, H. Xu, and J. Pang. Deximit: Learning bimanual dexterous manipulation from monocular human videos.arXiv preprint arXiv:2602.10105, 2026

  50. [58]

    Katharopoulos, A

    A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret. Transformers are RNNs: Fast autore- gressive transformers with linear attention. InInternational Conference on Machine Learning (ICML), 2020

  51. [59]

    Lugaresi, J

    C. Lugaresi, J. Tang, H. Nash, C. McClanahan, E. Uboweja, M. Hays, F. Zhang, C.-L. Chang, M. G. Yong, J. Lee, W.-T. Chang, W. Hua, M. Georg, and M. Grundmann. MediaPipe: A framework for building perception pipelines.arXiv preprint arXiv:1906.08172, 2019

  52. [60]

    Romero, D

    J. Romero, D. Tzionas, and M. J. Black. Embodied hands: Modeling and capturing hands and bodies together. InACM Transactions on Graphics (SIGGRAPH Asia), 2017

  53. [61]

    the value of the closest key

    O. Sim ´eoni, H. V . V o, M. Seitzer, F. Baldassarre, M. Oquab, C. Jose, V . Khalidov, M. Szafraniec, S. Yi, M. Ramamonjisoa, F. Massa, D. Haziza, L. Wehrstedt, J. Wang, T. Darcet, T. Moutakanni, L. Sentana, C. Roberts, A. Vedaldi, J. Tolan, J. Brandt, C. Couprie, J. Mairal, H...

Pith tools

Reviewed July 13, 2026 · model on record in the stance chip above.