REVIEW 3 major objections 6 minor 1 cited by
WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time
T0 review · 3 major / 6 minor · reviewed 2026-07-13 · grok-4.5
Pith's one-line read Unlabeled human play videos can steer a frozen robot foundation model at test time by adapting only a lightweight video-side memory.
desk verdict Solid systems paper: TTT memory for frozen WAMs from action-free human video is real and useful, but the big New-household numbers use in-scene human demos, so the pure “watch play elsewhere, steer here” story is softer than the abstract sells. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
WAM-TTT: residual TTT (test-time training) layers on the video expert of a frozen world-action model. Fast weights absorb human videos via video prediction plus key–value memory reconstruction; robot Queries read the adapted memory as a residual that steers action generation through shared visual-action dynamics.
What would settle it
On the same nine real-robot tasks and unseen household setting, replace the test-time fast-weight update with pure in-context conditioning on the identical human videos (or remove meta-training / the key–value loss) and check whether average progress collapses back toward the reported 7.1% baseline instead of remaining near 46%.
Extended reading notes
Core claim
A world-action model can be steered at test time from action-free human videos alone by updating only a lightweight fast-weight memory on the video expert, provided that memory was first meta-trained with paired human–robot data and a key–value reconstruction loss so that human Keys/Values become control-useful residuals for robot Queries.
Load-bearing premise
That phase-aligned paired human–robot meta-training produces a memory interface that stays useful for control when the only later updates come from unlabeled human videos on new tasks and scenes.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. WAM-TTT proposes a test-time training method to steer a frozen world-action model (built on LDA) using unlabeled human videos. A meta-training stage on phase-aligned paired human–robot data attaches video-side TTT residual branches and trains slow projections plus a key–value memory reconstruction loss so that human Keys/Values become a control-useful fast-weight memory; at deployment only those fast weights are updated by human-side video prediction and L_KVM while the WAM and action expert stay frozen. Real-robot evaluation on three embodiments and nine manipulation tasks reports large average gains in unseen household (New) settings over in-context human-video conditioning (WAM-ICL: 7.1% → 46.2%), the frozen backbone, co-training, and reimplemented EGOSCALE/π0.5 baselines, with ablations isolating meta-training, L_KVM, and TTT, plus data-ratio and pseudo-action studies.
Significance. If the result holds under a cleanly stated protocol, the paper offers a practical interface for RFM steering: absorb raw human play into a lightweight residual memory without robot actions, retargeting, or full fine-tuning, while keeping the foundation model frozen. The real-robot suite (3 embodiments, 9 tasks, Orig./New splits), the direct WAM-ICL control with the same human videos, and the informative ablations (especially pseudo-action harm and data-ratio iso-budget) are genuine strengths. The linear-attention witness in Appendix A is used only as motivation for L_KVM, not as a circular proof of performance. The work is a solid systems/methods contribution for world-action models and human-video transfer, contingent on clarifying how much of the New gain depends on scene-matched human videos.
major comments (3)
- [§3.3, §4.1–4.2, App. B, App. E.2, Table 1/C.1] Appendix B and Figures B.1–B.2 state that paired human demonstrations (and the human videos used for test-time TTT) are recorded with a GoPro “directly in the actual household environments that we later evaluate as the New setting,” while robot data is cubicle-only. Section 3.3 then adapts fast weights on those in-scene videos via L_vg + λ L_KVM. Table 1 / Table C.1 New numbers therefore compare TTT vs ICL under human videos that already share New lighting, clutter, and object instances—not pure transfer from out-of-scene human play. Appendix E.2’s “no in-scene human data” lab results are only qualitative. This is load-bearing for the abstract/intro framing of steering into new homes by watching human play. Please either (i) report quantitative New progress with human videos recorded outside the evaluation scene (or with cubicle-only human videos), or (ii) reframe claims and contribution
- [§4.2, Table 1, Table C.1] Table 1 New average (46.2%) is driven by large wins on several tasks, but Stamp Paper is a clear failure (WAM-TTT 8.3 vs LDA 33.3). The text attributes this to tight stamp geometry and household perturbation, yet the paper still claims consistent outperformance “across diverse manipulation tasks.” Either provide a failure analysis (e.g., whether human videos lack the corrective cue, or L_KVM overwrites a useful prior) or qualify the consistency claim and discuss when human-video TTT can hurt relative to the frozen backbone.
- [§4.3, Table 2] Table 2 ablations (meta-training, memory recon., TTT, LoRA) use only two tasks and 10 trials per cell, while the main claim rests on nine tasks × 25 trials. Given that w/o Meta Training collapses on Swap Place (0.0) and WAM-LoRA is 0.0 there, the design isolation is important but under-powered. Extend the protocol ablation to at least the full New suite (or a larger fixed subset) with the same 25-trial protocol as Table 1, or report confidence intervals so the component contributions are not over-read from two tasks.
minor comments (6)
- [§4.2, Table 1, Table C.1] No error bars or trial-level variance are reported for Table 1/C.1 despite 25 trials; adding mean±std or bootstrap intervals would make the +39.1 pt ICL gap easier to assess.
- [§3.2–3.3, Table B.1] Eq. (3)–(5) and Table B.1: N=1 inner SGD step is aggressive; a short sensitivity note on N and η_test would help readers judge stability of the fast-weight update.
- [Title, Figure 1, Abstract] Figure 1 / title use “W AM” / “WAM-TTT” spacing inconsistently; unify notation (WAM vs W AM) throughout.
- [App. D] Appendix D Stamp Paper rubric has a typo: “stamp s[uccessfully grasped.”
- [§2, §4.2] Related work on TTT and human-video transfer is thorough; a one-sentence contrast with MimicDroid [18] in the main text (beyond the citation list) would clarify the ICL baseline choice.
- [§4.4, Table 3] Table 3 generalization results are only on Deliver Drink; stating that scope in the caption would avoid over-generalizing “all perturbation types.”
Circularity Check
No circular derivation: empirical TTT method whose control claims rest on held-out robot progress, not on restating fitted inputs or self-citation uniqueness.
full rationale
WAM-TTT is a methods paper: meta-train a TTT branch on paired human–robot data (outer L_robot_WAM + inner L_vg + λ L_KVM), then at test time update only fast weights from unlabeled human videos while freezing the WAM. Performance claims (e.g., 46.2% vs 7.1% New progress vs WAM-ICL) are measured on real-robot trials with external baselines (π0.5, EGOSCALE, LDA), not obtained by renaming a fit as a prediction. Appendix A’s linear-attention “witness” is a closed-form motivation for L_KVM in the linear special case; it does not force the empirical robot results by construction. Citations to LDA [23] and Spatial-TTT [52] supply the backbone and TTT layer form—normal scaffolding, not a load-bearing uniqueness theorem that forbids alternatives. Phase-alignment and in-scene human-video protocol issues (if any) are experimental confounds, not circular reductions of equations to their inputs. No self-definitional loop, fitted-input-as-prediction, or ansatz-smuggled uniqueness chain is present. Score 0 with empty steps is the honest finding.
Assumptions & free parameters
free parameters (5)
- memory reconstruction weight λ =
4e-2
- inner SGD steps N =
1
- inner learning rates η_meta / η_test =
0.1 / 0.01
- TTT head dim d and fast-weight hidden width =
48 / 128
- meta-training data mix (robot, human) per task =
(100, 100)
assumptions (5)
- domain assumption A residual fast-weight update on video tokens can steer action generation through joint video-action attention without updating the action expert.
- ad hoc to paper Nearest-phase synchronization of human and robot episodes yields a valid human–robot alignment signal for meta-training.
- ad hoc to paper Minimizing key–value reconstruction on human Keys/Values produces a memory that robot Queries can usefully read (linear-attention witness).
- domain assumption Self-supervised video prediction on unlabeled human videos is a sufficient test-time objective for control-useful memory when the Q/K/V interface was meta-trained.
- domain assumption Standard diffusion / flow-matching multitask losses from the LDA backbone are valid outer objectives for joint WAM + TTT training.
invented entities (2)
-
Video-side TTT residual branches as human skill memory inside a frozen WAM
-
Key–value memory reconstruction loss L_KVM for human–robot meta-alignment
Cite this review
Pith. "Pith review of WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time." pith.science (2026). https://pith.science/paper/6VH44RPU
@misc{pith2026260706988,
author = {Pith},
title = {Pith review of: WAM-TTT: Steering World-Action Models by Watching Human Play at Test Time},
year = {2026},
howpublished = {\url{https://pith.science/paper/6VH44RPU}},
note = {Machine review of arXiv:2607.06988}
}
read the original abstract
Steering robot foundation models (RFMs) toward new task variants or user-preferred behaviors remains challenging, often requiring additional robot demonstrations, task-specific fine-tuning, or long-context conditioning. We present WAM-TTT, a test-time training framework for steering world action models from raw human videos. Rather than treating human videos as trajectories to imitate, WAM-TTT absorbs them into a lightweight adaptive memory inside a frozen WAM through self-supervised video prediction. To make this memory useful for control, we introduce a meta-training stage that aligns human demonstrations with robot behaviors using paired human-robot data and a key--value memory reconstruction objective. At test time, only unlabeled human videos are required to adapt the memory, while the pretrained WAM remains frozen. This enables efficient and reusable steering without robot actions, human-side annotations, or task-specific fine-tuning, while preserving the generalization ability of the foundation model. Extensive experiments show that WAM-TTT consistently outperforms in-context human-video conditioning baselines across diverse manipulation tasks and generalization settings.
Figures
Figures from the paper (2 more)
Forward citations
Cited by 1 Pith paper
-
StellaVLA: In-Context Structured Demonstration for Generalizable Vision-Language-Action Models
A vision-language-action policy conditioned on one automatically structured demonstration (sub-goals plus verbalized 3D/2D motion) achieves top scores on LIBERO, LIBERO-Plus, and VLA-Arena without fine-tuning.
Reference graph
Works this paper leans on
-
[1]
Yu et al
T. Yu et al. One-shot imitation from observing humans via domain-adaptive meta-learning. In RSS, 2018
2018
-
[2]
S. Bahl, A. Gupta, and D. Pathak. Human-to-robot imitation in the wild. InRSS, 2022
2022
-
[3]
M. Xu, Z. Xu, C. Chi, M. Veloso, and S. Song. Xskill: Cross embodiment skill discovery. In CoRL, 2023
2023
-
[4]
Bharadhwaj, A
H. Bharadhwaj, A. Gupta, V . Kumar, and S. Tulsiani. Towards generalizable zero-shot manip- ulation via translating human interaction plans. InICRA, 2024
2024
-
[5]
Hansen et al
N. Hansen et al. Self-supervised policy adaptation during deployment. InICLR, 2021
2021
-
[6]
M. Xu, Z. Xu, C. C. Pan, X. Zhu, C. Tomei, Y . Shen, Z. Wu, S.-R. Chen, J. B. Tenenbaum, T. Lozano-Perez, and S. Song. Flow as the cross-domain manipulation interface. InCoRL, 2024
2024
- [7]
- [8]
Show all 61 references
-
[9]
Grauman et al
K. Grauman et al. Ego-exo4d: Understanding skilled human activity from first- and third- person perspectives. InCVPR, 2024
2024
-
[10]
Zheng et al
R. Zheng et al. Egoscale: Scaling dexterous manipulation with diverse egocentric human data. arXiv preprint arXiv:2602.16710, 2026
2026
-
[11]
Chen et al
H. Chen et al. Vidbot: Learning generalizable 3d actions from in-the-wild 2d human videos for zero-shot robotic manipulation. InCVPR, 2025
2025
-
[12]
Kim et al
H. Kim et al. Uniskill: Imitating human videos via cross-embodiment skill representations. In CoRL, 2025
2025
-
[13]
Z. Chen, S. Chen, E. Arlaud, I. Laptev, and C. Schmid. Vividex: Learning vision-based dex- terous manipulation from human videos. InICRA, 2025
2025
-
[14]
C. Wang, L. Fan, J. Sun, R. Zhang, L. Fei-Fei, D. Xu, Y . Zhu, and A. Anandkumar. Mimicplay: Long-horizon imitation learning by watching human play. InCoRL, 2023. arXiv:2302.12422
2023 arXiv
-
[15]
Bharadhwaj, A
H. Bharadhwaj, A. Gupta, S. Tulsiani, and V . Kumar. Zero-shot robot manipulation from passive human videos.arXiv preprint arXiv:2302.02011, 2023
2023 arXiv
-
[16]
V . Jain, M. Attarian, N. J. Joshi, A. Wahid, D. Driess, Q. Vuong, P. R. Sanketi, P. Ser- manet, S. Welker, C. Chan, I. Gilitschenski, Y . Bisk, and D. Dwibedi. Vid2robot: End- to-end video-conditioned policy learning with cross-attention transformers.arXiv preprint arXiv:2403...
2024 arXiv
-
[17]
Bharadhwaj, D
H. Bharadhwaj, D. Dwibedi, A. Gupta, S. Tulsiani, C. Doersch, T. Xiao, D. Shah, F. Xia, D. Sadigh, and S. Kirmani. Gen2act: Human video generation in novel scenarios enables generalizable robot manipulation.arXiv preprint arXiv:2409.16283, 2024
2024 arXiv
-
[18]
R. Shah, S. Liu, Q. Wang, Z. Jiang, S. Kumar, M. Seo, R. Mart ´ın-Mart´ın, and Y . Zhu. Mim- icdroid: In-context learning for humanoid robot manipulation from human play videos.arXiv preprint arXiv:2509.09769, 2025. 10
2025 arXiv
-
[19]
Chi, C.-K
X. Chi, C.-K. Fan, H. Zhang, X. Qi, R. Zhang, A. Chen, C.-m. Chan, W. Xue, Q. Liu, S. Zhang, et al. Eva: An embodied world model for future video anticipation.arXiv preprint arXiv:2410.15461, 2024
2024 arXiv
-
[20]
X. Chi, P. Jia, C.-K. Fan, X. Ju, W. Mi, K. Zhang, Z. Qin, W. Tian, K. Ge, H. Li, et al. Wow: Towards a world omniscient world model through embodied interaction.arXiv preprint arXiv:2509.22642, 2025
2025
-
[21]
Zhang, X
J. Zhang, X. Chen, A.-J. Chen, C. Lv, D. mei Li, G. Zhou, H. Yin, H. Yuan, H. Li, J. Li, J. Zhang, J. Zhou, K. Gao, K. Yan, L. Jiang, N. Tang, P. Lin, Q. Peng, S.-S. Yin, T. Wu, T. Yan, X. Xu, Y . Shu, Y . Zhang, Y . Wang, Y . Wang, Y . Chen, Y . Xu, Y . Huang, Y . Chen, Z. Zh...
2026
-
[22]
C. Zhu, R. Yu, S. Feng, B. Burchfiel, P. Shah, and A. Gupta. Unified world models: Cou- pling video and action diffusion for pretraining on large robotic datasets.arXiv preprint arXiv:2504.02792, 2025
2025 arXiv
-
[23]
J. Lyu, K. Liu, X. Zhang, H. Liao, Y . Feng, W. Zhu, T. Shen, J. Chen, J. Zhang, Y . Dong, W. Cui, S. Qi, S. Wang, Y . Zheng, M. Yan, X. Shi, H. Li, D. Zhao, M.-Y . Liu, Z. Zhang, L. Yi, Y . Wang, and H. Wang. LDA-1B: Scaling latent dynamics action model via universal embodied...
2026 arXiv
-
[24]
H. Bi, H. Tan, S. Xie, Z. Wang, S. Huang, H. Liu, R. Zhao, Y . Feng, C. Xiang, Y . Rong, et al. Motus: A unified latent action world model. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 35101–35113, 2026
2026
-
[25]
M. Team, C. Xiang, F. Bao, H. Liu, H. Tan, H. Bi, J. Li, J. Liu, J. Pang, K. Jing, et al. Mo- tubrain: An advanced world action model for robot control.arXiv preprint arXiv:2604.27792, 2026
2026 arXiv
-
[26]
Zhang, W
Y . Zhang, W. Zhang, Z. Qi, H. Zhang, H. Lin, J. Zhang, Y . Mu, X. Yang, W. Zeng, and X. Jin. Imagewam: Do world action models really need video generation, or just image editing?, 2026. URLhttps://arxiv.org/abs/2606.19531
2026 arXiv
-
[27]
H. Yu, H. Lin, J. Zhang, W. Zhang, C. Gu, H. Li, and P. Tan. Maskwam: Unifying mask prompting and prediction for world-action models.arXiv preprint arXiv:2606.13515, 2026
2026 arXiv
-
[28]
B. Peng, W. Zhang, L. Xu, Z. Qi, J. Zhang, H. Liu, W. Zeng, and X. Jin. Reworld: Multi- dimensional reward modeling for embodied world models.arXiv preprint arXiv:2601.12428, 2026
2026
-
[29]
S. Ye, Y . Ge, K. Zheng, S. Gao, S. Yu, G. Kurian, S. Indupuru, Y . L. Tan, C. Zhu, J. Xi- ang, A. Malik, K. Lee, W. Liang, N. Ranawaka, J. Gu, Y . Xu, G. Wang, F. Hu, A. Narayan, J. Bjorck, J. Wang, G. Kim, D. Niu, R. Zheng, Y . Xie, J. Wu, Q. Wang, R. Julian, D. Xu, Y . Du, ...
2026 arXiv
-
[30]
A. Ye, B. Wang, C. Ni, G. Huang, G. Zhao, H. Li, H. Li, J. Li, J. Lv, J. Liu, M. Cao, P. Li, Q. Deng, W. Mei, X. Wang, X. Chen, X. Zhou, Y . Wang, Y . Chang, Y . Li, Y . Zhou, Y . Ye, Z. Liu, and Z. Zhu. Gigaworld-policy: An efficient action-centered world-action model.arXiv p...
2026
-
[31]
L. Li, Q. Zhang, Y . Luo, S. Yang, R. Wang, F. Han, M. Yu, Z. Gao, N. Xue, X. Zhu, Y . Shen, and Y . Xu. Causal world modeling for robot control.arXiv preprint arXiv:2601.21998, 2026. 11
2026 arXiv
-
[32]
T. Yuan, Z. Dong, Y . Liu, and H. Zhao. Fast-wam: Do world action models need test-time future imagination?arXiv preprint arXiv:2603.16666, 2026. URLhttps://arxiv.org/ abs/2603.16666
2026 arXiv
-
[33]
J. Cai, L. Ling, S. Chu, Z. Liu, J. Kang, Z. Liang, W. Xu, Y . Mao, W. Zhang, X. Yang, R. Ying, R. Zheng, and Y . Mu. Aha-wam: Asynchronous horizon-adaptive world-action modeling with observation-guided context routing.arXiv preprint arXiv:2606.09811, 2026
2026 arXiv
-
[34]
Q. Feng, J. Yu, J. Liu, Y . Jia, Z. Wu, H. Chen, Z. Qian, S. Gu, P. Jia, S. Ma, and S. Zhang. Harmowam: Harmonizing generalizable and precise manipulation via adaptive world action models, 2026
2026
-
[35]
H. Luo, W. Zhang, Y . Feng, S. Zheng, H. Xu, C. Xu, Z. Xi, Y . Fu, and Z. Lu. Being-h0. 7: A latent world-action model from egocentric videos.arXiv preprint arXiv:2605.00078, 2026
2026 arXiv
-
[36]
J. Guo, Q. Li, P. Li, Z. Chen, N. Sun, Y . Su, H. Wang, Y . Zhang, X. Li, and H. Liu. Unified 4d world action modeling from video priors with asynchronous denoising.arXiv preprint arXiv:2604.26694, 2026
2026 arXiv
-
[37]
J. Lyu, Z. Li, X. Shi, C. Xu, Y . Wang, and H. Wang. Dywa: Dynamics-adaptive world action model for generalizable non-prehensile manipulation.arXiv preprint arXiv:2503.16806, 2025
2025 arXiv
-
[38]
Agarwal, A
N. Agarwal, A. Ali, J. Allen, M. Antolini, A. Aubame, A. Azzolini, J. Bai, M. Bala, Y . Bal- aji, J. Bapst, et al. Cosmos 3: Omnimodal world models for physical ai.arXiv preprint arXiv:2606.02800, 2026
2026 arXiv
-
[39]
T. Ma, J. Zheng, Z. Wang, C. Jiang, A. Cui, J. Liang, and S. Yang. Dit4dit: Jointly modeling video dynamics and actions for generalizable robot control.arXiv preprint arXiv:2603.10448, 2026
2026
-
[40]
Physical Intelligence, B. Ai, A. Amin, R. J. Aniceto, A. Balakrishna, G. Balke, K. Black, G. Bokinsky, S. Cao, T. Charbonnier, et al.π 0.7: A steerable generalist robotic foundation model with emergent capabilities.arXiv preprint arXiv:2604.15483, 2026
2026 arXiv
-
[41]
Liu et al
Y . Liu et al. Oa-wam: Object-addressable world action model for robust robot manipulation. 2026
2026
-
[42]
Y . Yang, S. Zeng, T. Lin, X. Chang, D. Qi, J. Xiao, H. Liu, R. Chen, Y . Chen, D. Huo, et al. Abot-m0: Vla foundation model for robotic manipulation with action manifold learning.arXiv preprint arXiv:2602.11236, 2026
2026 arXiv
-
[43]
R. Chen, Y . Yang, Z. Tang, D. Huo, T. Lin, H. Wu, H. Liu, Y . Chen, L. Zheng, B. Yuan, T. Li, M. Wang, D. Qi, B. Hu, W. Mei, Y . Xuan, H. Yang, Y . Zhu, M. Xu, Z. Ma, and X. Chang. Abot-m0.5: Unified mobility-and-manipulation world action model.arXiv preprint arXiv:2607.00678, 2026
2026 arXiv
-
[44]
M. J. Kim, Y . Gao, T.-Y . Lin, Y .-C. Lin, Y . Ge, G. Lam, P. Liang, S. Song, M.-Y . Liu, C. Finn, and J. Gu. Cosmos policy: Fine-tuning video models for visuomotor control and planning. In International Conference on Learning Representations (ICLR), 2026
2026
-
[45]
Y . Hu, Y . Guo, P. Wang, X. Chen, Y .-J. Wang, J. Zhang, K. Sreenath, C. Lu, and J. Chen. Video prediction policy: A generalist robot policy with predictive visual representations.arXiv preprint arXiv:2412.14803, 2024
2024 arXiv
-
[46]
Y . Sun, X. Wang, Z. Liu, J. Miller, A. A. Efros, and M. Hardt. Test-time training with self- supervision for generalization under distribution shifts. InICML, 2020
2020
-
[47]
D. Wang, E. Shelhamer, S. Liu, B. Olshausen, and T. Darrell. Tent: Fully test-time adaptation by entropy minimization. InICLR, 2021. 12
2021
-
[48]
Y . Liu, P. Kothari, B. van Delft, B. Bellot-Gurlet, T. Mordan, and A. Alahi. Ttt++: When does self-supervised test-time training fail or thrive? InNeurIPS, 2021
2021
-
[49]
Gandelsman, Y
Y . Gandelsman, Y . Sun, X. Chen, and A. A. Efros. Test-time training with masked autoen- coders. InNeurIPS, 2022
2022
-
[50]
Y . Sun, X. Li, K. Dalal, J. Xu, A. Vikram, G. Zhang, Y . Dubois, X. Chen, X. Wang, S. Koyejo, T. Hashimoto, and C. Guestrin. Learning to (learn at test time): Rnns with expressive hidden states. InICML, 2025
2025
-
[51]
Behrouz, P
A. Behrouz, P. Zhong, and V . Mirrokni. Titans: Learning to memorize at test time.arXiv preprint arXiv:2501.00663, 2025
2025 arXiv
-
[52]
F. Liu, D. Wu, J. Chi, Y . Cai, Y .-H. Hung, X. Yu, H. Li, H. Hu, Y . Rao, and Y . Duan. Spatial-ttt: Streaming visual-based spatial intelligence with test-time training.arXiv preprint arXiv:2603.12255, 2026
2026
-
[53]
S. Yang, Y . Ze, and H. Xu. Movie: Visual model-based policy adaptation for view generaliza- tion. InNeurIPS, 2023
2023
-
[54]
Z. Bai, C. Gao, and M. Z. Shou. Evolve-vla: Test-time training from environment feedback for vision-language-action models.arXiv preprint arXiv:2512.14666, 2025
2025
-
[55]
C. Liu, Y . Liu, T. Wang, Q. Zhuang, J. C. Liang, W. Yang, R. Xu, Q. Wang, D. Liu, and C. Han. On-the-fly vla adaptation via test-time reinforcement learning.arXiv preprint arXiv:2601.06748, 2026
2026 arXiv
-
[56]
Black, N
K. Black, N. Brown, D. Driess, A. Esmail, M. Equi, C. Finn, N. Fusai, L. Groom, K. Haus- man, B. Ichter, S. Jakubczak, T. Jones, L. Ke, S. Levine, A. Li-Bell, M. Mothukuri, S. Nair, K. Pertsch, L. X. Shi, L. Smith, J. T. Springenberg, K. Stachowicz, J. Tanner, Q. Vuong, H. Wal...
2025 arXiv
-
[57]
J. Mu, S. Yang, Y . Bao, H. Bae, T. Wei, L. Xu, B. Li, H. Xu, and J. Pang. Deximit: Learning bimanual dexterous manipulation from monocular human videos.arXiv preprint arXiv:2602.10105, 2026
2026
-
[58]
Katharopoulos, A
A. Katharopoulos, A. Vyas, N. Pappas, and F. Fleuret. Transformers are RNNs: Fast autore- gressive transformers with linear attention. InInternational Conference on Machine Learning (ICML), 2020
2020
-
[59]
Lugaresi, J
C. Lugaresi, J. Tang, H. Nash, C. McClanahan, E. Uboweja, M. Hays, F. Zhang, C.-L. Chang, M. G. Yong, J. Lee, W.-T. Chang, W. Hua, M. Georg, and M. Grundmann. MediaPipe: A framework for building perception pipelines.arXiv preprint arXiv:1906.08172, 2019
1906 arXiv
-
[60]
Romero, D
J. Romero, D. Tzionas, and M. J. Black. Embodied hands: Modeling and capturing hands and bodies together. InACM Transactions on Graphics (SIGGRAPH Asia), 2017
2017
-
[61]
the value of the closest key
O. Sim ´eoni, H. V . V o, M. Seitzer, F. Baldassarre, M. Oquab, C. Jose, V . Khalidov, M. Szafraniec, S. Yi, M. Ramamonjisoa, F. Massa, D. Haziza, L. Wehrstedt, J. Wang, T. Darcet, T. Moutakanni, L. Sentana, C. Roberts, A. Vedaldi, J. Tolan, J. Brandt, C. Couprie, J. Mairal, H...
2025 arXiv
Reviewed July 13, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.