REVIEW 3 major objections 4 minor 86 references
RoboTTT claims that scaling a robot policy's pretraining context to 8K timesteps yields steady closed-loop performance gains and unlocks one-shot imitation and on-the-fly recovery.
Reviewed by Pith at T0; open to challenge. T0 means a machine referee read the full paper against a public rubric. the ladder, T0–T4 →
T0 review · deepseek-v4-flash
2026-08-01 23:38 UTC pith:NYS7BWZ2
load-bearing objection Solid systems paper; the scaling-axis headline is underdetermined by a compute/curriculum confound between the 1K and 8K pretraining runs, and the trial counts are thin. the 3 major comments →
RoboTTT: Context Scaling for Robot Policies
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
Core claim
RoboTTT is a sequence model for robot control built by adding test-time-training (TTT) layers to a pretrained vision-language-action flow-matching policy. Its recurrent state is a set of fast weights: a small MLP updated by gradient descent on an inner self-supervised loss at every timestep, during training and at deployment, so the history is compressed into weight space and later retrieved when producing actions. The training recipe — independent noise levels per action chunk in the sequence (sequence action forcing) plus truncated backpropagation through time — makes training on 8K-timestep sequences feasible under a fixed GPU memory budget. On three real-robot assembly tasks the paper re
What carries the argument
The load-bearing mechanism is the TTT layer with fast weights: a two-layer MLP parameterized by weights updated by gradient descent on an inner loss and then applied to the query, so the model's memory is the weight state itself rather than a fixed-size vector or cached keys and values. Register tokens carry vision-language information across timesteps, a learned tanh gate protects the pretrained behavior early in training, and sequence action forcing with TBPTT lets context grow without memory blow-up. DAgger Distillation is a secondary mechanism: failures update fast weights while the loss is masked to human corrections, distilling the failure-to-correction mapping into the weight state.
Load-bearing premise
The headline comparisons rest on 10–20 rollouts per condition with no reported confidence intervals or multiple seeds, so the observed 57–63% gaps and monotone scaling curve could in principle be sampling noise.
What would settle it
Rerun the main comparisons (8K vs 1K context, and the 128-to-8K scaling series) with at least 50 rollouts per condition and per-configuration confidence intervals, ideally with blinded rubric scoring. If the 8K advantage over 1K falls below the noise floor, or the scaling curve flattens or reverses on a third task, the paper's central claim that pretraining context length yields steady closed-loop gains is falsified.
If this is right
- Pretraining context length can be treated as a scaling axis: longer context translates into higher closed-loop task completion for the same model, with no observed saturation between 128 and 8K timesteps.
- Long-context conditioning is sufficient to enable one-shot imitation from a single human video of an unseen task configuration, a capability short-context and recurrent-memory baselines lack in this setup.
- A policy can learn to improve on the fly: distilling failure-to-correction pairs into fast weights yields better recovery than standard DAgger fine-tuning on corrections alone.
- Because the recurrent state is fixed-size (fast weights), inference cost stays constant in context length, making multi-minute memories practical at 30 Hz control.
- The scaling benefit is tied to the update rule: a gradient-descent fast model benefits from longer context, while a gated linear recurrent memory does not.
Where Pith is reading between the lines
- Editorial inference: if the scaling trend is real, the gains may persist beyond 8K until the fast-weight MLP's capacity or the meta-learned update dynamics saturate; a natural next test is measuring task completion at 16K and 32K with matched compute.
- Editorial inference: the failure-as-context principle behind DAgger Distillation could extend to other data sources — such as preference judgments or evaluator critiques — where suboptimal behavior is available alongside corrections, without needing new imitation targets.
- Editorial inference: the one-shot video-imitation result suggests a route to task specification that does not rely on language; a testable extension is whether conditioning on multiple videos or videos with distractors improves robustness and whether fast weights can be reset between tasks.
- Editorial inference: if context length is a scaling axis, the compute and training-cost tradeoff becomes central; cheaper TTT training techniques or chunkwise training would determine whether this axis is practical at frontier scale.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper introduces RoboTTT, a Test-Time-Training robot policy built on GR00T N1.7, with fast-weight recurrent states updated by gradient descent at train and test time. The training recipe combines sequence action forcing with truncated backpropagation through time to scale pretraining context to 8K timesteps at fixed inference cost. On three real-robot bimanual assembly tasks, the authors report that RoboTTT-8K outperforms a single-step baseline and a matched Gated DeltaNet recurrent memory baseline, that closed-loop performance rises steadily as pretraining context length grows from 128 to 8K, and that long-context conditioning enables one-shot imitation from human video, on-the-fly recovery via DAgger Distillation, and perturbation robustness.
Significance. If the empirical claims hold, the paper is significant: it would be the first demonstration that pretraining context length is a scaling axis for closed-loop robot manipulation, and it would show that gradient-based fast weights are a practical sequence-memory mechanism for deployed VLA policies. The evaluation has real strengths: held-out Circuit configurations (20 train / 60 test), matched post-training across methods, a GDN recurrent baseline with matched layer placement and parameter count, ablations of action forcing and fast-model expressivity, and unusually detailed rubrics and implementation notes in the appendices. The central scaling claim, however, is currently underdetermined because context length is not varied with matched compute or matched curriculum, and the statistical evidence is thin (10-20 rollouts per condition, no confidence intervals). These issues are fixable and do not undermine the architectural contribution itself, but they block acceptance of the paper as it stands.
major comments (3)
- [Fig. 8; §3.4; §A.2] The central claim that 'context length is a new scaling axis' is not isolated from pretraining compute. §A.2 states that per-device batch is 4 (global 64) for context lengths ≤4K and 1 (global 16) for >4K, and §3.4 says context length is 'gradually increased' to the target. Thus the 8K run consumes 16×8192=131K timesteps per optimization step versus 64×1024=64K for the 1K run, a 2× gap, and the curriculum also differs. 'The same model pretrained with 1K-timestep context' is therefore not a controlled comparison. GDN's flat curve weakens a pure-compute explanation, but GDN is a different model class, so it cannot control for an interaction between compute and RoboTTT's meta-learned fast-weight dynamics. Please add a compute-matched control (e.g., a 1K model trained with the same global batch/curriculum and scaled total FLOPs) or report token throughput and a matched-batch schedule for eve
- [§4; Fig. 8; Tables 1-3] All headline comparisons rest on 10-20 rollouts per condition with no confidence intervals, per-configuration variance, or multiple seeds. Examples: Gear Bot full successes are 2/10 vs 0/10, one-shot imitation is 6/10 vs 0/10 (Table 2), and Pup Go Car is 9/20 vs 3/20 (Table 1). Some differences may be genuine, but as reported they are not statistically quantified, and the 63%/57% scaling gaps in Fig. 8 use a single average without error bars. Please provide bootstrap confidence intervals over held-out configurations (or over seeds), and report per-configuration outcomes. This is load-bearing for the scaling and one-shot-imitation claims.
- [§3.4 vs §4; one-shot imitation] The paper's context-length protocol is internally ambiguous. §3.4 says all models are post-trained on each downstream task at 1K context length, and §4 says sequence models 'use a 1K-timestep context.' Yet one-shot imitation is trained by concatenating a human video and a robot trajectory into a single training sequence, which can easily exceed 1K timesteps for 30 Hz control. Either the one-shot experiments use a longer post-training context (contradicting §3.4) or the effective training context is only 1K (contradicting the '8K context' framing). Please clarify the exact context used for post-training and evaluation in every experiment, and state how 'RoboTTT-8K' should be interpreted if deployment is at a shorter context.
minor comments (4)
- [Fig. 8 caption] 'All evaluations in this figure predate the DAgger training used for Pup Go Car in the main results' is ambiguous: does it mean the Pup Go Car component of the figure was run before DAgger was introduced, and if so, why is the DAgger-trained model the one reported in Table 1? Please clarify the relationship between Fig. 8 and the main evaluation.
- [Fig. 12] The ablation figure reports only relative improvements without numeric values or error bars. Please add exact values, confidence intervals, and trial counts.
- [Table 3 and text] The text says 'the long-context methods react successfully more often' but in the tire condition RoboTTT and GDN both recover 18/20; the claim of an advantage over GDN is supported only by the roof condition. Please soften or qualify.
- [§4, baseline description] GR00T N1.7 Hist. is described as using 'one history frame'; the claim that 'history alone does not reliably help' is based on a single additional frame and on Pup Go Car, where the history baseline is worse than no history. This is a very narrow form of history augmentation; please acknowledge the limitation.
Circularity Check
No significant circularity: the paper's claims are empirical comparisons against external task rubrics; no fitted parameter is relabeled as a prediction.
full rationale
This is an empirical systems paper, not a derivation chain. The central claim—that scaling pretraining context length from 128 to 8K timesteps yields steady closed-loop gains (Sec. 4, Fig. 8)—rests on comparing the same RoboTTT architecture pretrained at different context lengths and post-trained on identical task data with rubric-based completion scores. No model parameter is fitted to the reported completion percentages and then renamed as a prediction; the 8K-vs-1K comparison is an internal experimental contrast, not an identity by construction. The TTT update/apply equations (Eqs. 1–2), the sequence flow-matching loss (Eq. 5), tanh gating (Eq. 3), and TBPTT are stated model mechanics rather than conclusions derived from the benchmarks. Self-citations to GR00T N1.7 and Egoscale are prior-art backbone and pretraining-data usage; these are externally released resources and do not load-bear the scaling conclusion. The one-shot imitation and DAgger Distillation results are trained capabilities evaluated on held-out trials, and while their generalization could be debated, they do not reduce to the training inputs by definition. The main weaknesses—confounding context length with pretraining compute/batch schedule and small trial counts without confidence intervals—are correctness and experimental-validity concerns, not circularity. Accordingly, no circular step can be exhibited, and the score is 0.
Axiom & Free-Parameter Ledger
free parameters (6)
- Inner fast-weight learning rate η =
learned (base 0.1)
- Tanh gate initialization α =
0.001
- Register token count N =
16
- Fast model capacity =
two-layer MLP, ~10M params per TTT layer
- Sequence action forcing noise schedule =
τ=s(1-u), u~Beta(1.5,1), s=0.999
- Pretraining/post-training context lengths =
up to 8K / 1K timesteps
axioms (7)
- domain assumption The real-robot evaluation protocol (fixed initial placements, 4 cameras, 30Hz control) yields measurements that transfer to claims about robot policies in general.
- domain assumption Rubric-based task completion scores are valid partial-credit measures of assembly progress.
- domain assumption The pretraining mixture (tabletop bimanual robot + egocentric human data) is representative enough for TTT fast-weight dynamics learned there to transfer to the three downstream tasks.
- domain assumption Fast-weight updates by gradient descent remain stable and informative across 8K timesteps during deployment.
- domain assumption GDN is a fair, sufficiently strong recurrent-memory baseline; RoboTTT-vs-GDN differences are attributed to the update rule.
- domain assumption The 10-20 rollouts per condition are enough support for the reported success counts and scaling curve.
- domain assumption In the one-shot imitation experiment, the human video is the only cue identifying the target configuration.
read the original abstract
Recent robot foundation models operate with single-step or short-history visuomotor context. We introduce Test-Time-Training Robot Policies (RoboTTT), a robot model and training recipe that scale visuomotor context to 8K timesteps, three orders of magnitude beyond state-of-the-art policies, without growing inference latency. At this context length, we unlock new robot capabilities: one-shot in-context imitation from human video demonstrations, on-the-fly policy improvement, robustness to perturbations, and stronger performance on multi-stage, long-horizon tasks. We also observe, for the first time, steady gains in closed-loop performance as pretraining context length scales. At its core, RoboTTT integrates Test-Time Training into robot foundation models such as Vision-Language-Action policies, yielding a sequence model whose recurrent state consists of fast weights, parameters updated by gradient descent during both training and inference, compressing histories into weight space and retrieving contextual information for long-context conditioning. To scale training context length, the recipe combines sequence action forcing with truncated backpropagation through time. On challenging real-robot manipulation tasks, RoboTTT improves overall performance by 87% over the single-step context baseline and fully completes a five-minute, ten-stage assembly task, which no baseline ever does. RoboTTT trained with 8K-timestep context outperforms the same model pretrained with 1K timesteps by 62%, suggesting context length as a new scaling axis for robot foundation models. Videos are available at https://research.nvidia.com/labs/gear/robottt/
Figures
Reference graph
Works this paper leans on
-
[1]
Flamingo: a visual language model for few-shot learning.arXiv preprint arXiv: 2204.14198, 2022
Jean-Baptiste Alayrac, Jeff Donahue, Pauline Luc, Antoine Miech, Iain Barr, Yana Hasson, Karel Lenc, Arthur Mensch, Katie Millican, Malcolm Reynolds, Roman Ring, Eliza Rutherford, Serkan Cabi, Tengda Han, Zhitao Gong, Sina Samangooei, Marianne Monteiro, Jacob Menick, Sebastian Borgeaud, Andrew Brock, Aida Nematzadeh, Sahand Sharifzadeh, Mikolaj Binkowski,...
Pith/arXiv arXiv 2022
-
[2]
Zechen Bai, Chen Gao, and Mike Zheng Shou. Evolve-vla: Test-time training from environment feedback for vision-language-action models.arXiv preprint arXiv: 2512.14666, 2025. 12
arXiv 2025
-
[3]
Titans: Learning to memorize at test time.arXiv preprint arXiv: 2501.00663, 2024
Ali Behrouz, Peilin Zhong, and Vahab Mirrokni. Titans: Learning to memorize at test time.arXiv preprint arXiv: 2501.00663, 2024. 12
Pith/arXiv arXiv 2024
-
[4]
Ali Behrouz, Meisam Razaviyayn, Peilin Zhong, and Vahab Mirrokni. It’s all connected: A journey through test-time memorization, attentional bias, retention, and online optimization.arXiv preprint arXiv: 2504.13173, 2025. 12
Pith/arXiv arXiv 2025
-
[5]
Kevin Black, Noah Brown, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Lachy Groom, Karol Hausman, Brian Ichter, et al.𝜋0: A vision-language-action flow model for general robot control, 2024.URL https://arxiv. org/abs/2410.24164, 2024. 2, 11, 12
Pith/arXiv arXiv 2024
-
[6]
Lee, Maria Bauzá Villalonga, Todor Davchev, Yuxiang Zhou, Agrim Gupta, A
Konstantinos Bousmalis, Giulia Vezzani, Dushyant Rao, Coline Devin, Alex X. Lee, Maria Bauzá Villalonga, Todor Davchev, Yuxiang Zhou, Agrim Gupta, A. Raju, Antoine Laurens, Claudio Fantacci, Valentin Dalibard, Martina Zambelli, M. F. Martins, Rugile Pevceviciute, M. Blokzijl, Misha Denil, Nathan Batchelor, Thomas Lampe, Emilio Parisotto, Konrad Zolna, Sco...
-
[7]
Irpan, Tomas Jackson, Sally Jesmonth, Nikhil J
AnthonyBrohan, NoahBrown, JusticeCarbajal, YevgenChebotar, JosephDabis, ChelseaFinn, K.Gopalakr- ishnan, Karol Hausman, Alexander Herzog, Jasmine Hsu, Julian Ibarz, Brian Ichter, A. Irpan, Tomas Jackson, Sally Jesmonth, Nikhil J. Joshi, Ryan C. Julian, Dmitry Kalashnikov, Yuheng Kuang, Isabel Leal, Kuang-Huei Lee, S. Levine, Yao Lu, U. Malla, D. Manjunath...
-
[8]
Choromanski, Tianli Ding, Danny Driess, Kumar Avinava Dubey, Chelsea Finn, Peter R
Anthony Brohan, Noah Brown, Justice Carbajal, Yevgen Chebotar, K. Choromanski, Tianli Ding, Danny Driess, Kumar Avinava Dubey, Chelsea Finn, Peter R. Florence, Chuyuan Fu, Montse Gonzalez Arenas, K. Gopalakrishnan, Kehang Han, Karol Hausman, Alexander Herzog, Jasmine Hsu, Brian Ichter, A. Irpan, 13 RoboTTT: Context Scaling for Robot Policies Nikhil J. Jos...
-
[9]
Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, Arvind Neelakantan,PranavShyam,GirishSastry,AmandaAskell,SandhiniAgarwal,ArielHerbert-Voss,Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel M. Ziegler, Jeffrey Wu, Clemens Winter, Christopher Hesse, Mark Chen, Eric Sigler, Mateusz Litwin, Scott Gra...
Pith/arXiv arXiv 2005
-
[10]
Boyuan Chen, Diego Marti Monso, Yilun Du, Max Simchowitz, Russ Tedrake, and Vincent Sitzmann. Diffusion forcing: Next-token prediction meets full-sequence diffusion.arXiv preprint arXiv: 2407.01392,
-
[11]
Ttt3r: 3d reconstruction as test-time training, 2026
Xingyu Chen, Yue Chen, Yuliang Xiu, Andreas Geiger, and Anpei Chen. Ttt3r: 3d reconstruction as test-time training, 2026. URLhttps://arxiv.org/abs/2509.26645. 12
Pith/arXiv arXiv 2026
-
[12]
Robomme: Benchmarking and understanding memory for robotic generalist policies
Yinpei Dai, Hongze Fu, Jayjun Lee, Yuejiang Liu, Haoran Zhang, Jianing Yang, Chelsea Finn, Nima Fazeli, and Joyce Chai. Robomme: Benchmarking and understanding memory for robotic generalist policies. arXiv preprint arXiv:2603.04639, 2026. 2
Pith/arXiv arXiv 2026
-
[13]
One-minute video generation with test-time training
Karan Dalal, Daniel Koceja, Jiarui Xu, Yue Zhao, Shihao Han, Ka Chun Cheung, Jan Kautz, Yejin Choi, Yu Sun, and Xiaolong Wang. One-minute video generation with test-time training. InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 17702–17711, June 2025. 12
2025
-
[14]
Vision transformers need registers
Timothée Darcet, Maxime Oquab, Julien Mairal, and Piotr Bojanowski. Vision transformers need registers. InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11,
2024
-
[15]
Causal confusion in imitation learning
Pim de Haan, Dinesh Jayaraman, and Sergey Levine. Causal confusion in imitation learning. In H. Wal- lach, H. Larochelle, A. Beygelzimer, F. d'Alché-Buc, E. Fox, and R. Garnett, editors,Advances in Neural Information Processing Systems, volume 32. Curran Associates, Inc., 2019. URLhttps://proceedings. neurips.cc/paper_files/paper/2019/file/947018640bf36a2...
arXiv 2019
-
[16]
Longrope: Extending LLM context window beyond 2 million tokens
Yiran Ding, Li Lyna Zhang, Chengruidong Zhang, Yuanyuan Xu, Ning Shang, Jiahang Xu, Fan Yang, and Mao Yang. Longrope: Extending LLM context window beyond 2 million tokens. In Ruslan Salakhutdinov, Zico Kolter, Katherine A. Heller, Adrian Weller, Nuria Oliver, Jonathan Scarlett, and Felix Berkenkamp, editors,Forty-first International Conference on Machine ...
2024
-
[17]
Knowledge insulating vision-language-action models: Train fast, run fast, generalize better.Advances in Neural Information Processing Systems, 38:102867–102888,
Danny Driess, Jost Springenberg, Brian Ichter, Lili Yu, Adrian Li-Bell, Karl Pertsch, Allen Ren, Homer Walke, Quan Vuong, Lucy Xiaoyang Shi, et al. Knowledge insulating vision-language-action models: Train fast, run fast, generalize better.Advances in Neural Information Processing Systems, 38:102867–102888,
-
[18]
One-shot imitation learning.Advances in neural information processing systems, 30, 2017
Yan Duan, Marcin Andrychowicz, Bradly Stadie, OpenAI Jonathan Ho, Jonas Schneider, Ilya Sutskever, Pieter Abbeel, and Wojciech Zaremba. One-shot imitation learning.Advances in neural information processing systems, 30, 2017. 2, 11 14 RoboTTT: Context Scaling for Robot Policies
2017
-
[19]
Molmoact2: Action reasoning models for real-world deployment, 2026
Haoquan Fang, Jiafei Duan, Donovan Clay, Sam Wang, Shuo Liu, Weikai Huang, Xiang Fan, Wei-Chuan Tsai, Shirui Chen, Yi Ru Wang, Shanli Xing, Jaemin Cho, Jae Sung Park, Ainaz Eftekhar, Peter Sushko, Karen Farley, Angad Wadhwa, Cole Harrison, Winson Han, Ying-Chun Lee, Eli VanderBilt, Rose Hendrix, Suveen Ellawela, Lucas Ngoo, Joyce Chai, Zhongzheng Ren, Ali...
Pith/arXiv arXiv 2026
-
[20]
In-place test-time training
Guhao Feng, Shengjie Luo, Kai Hua, Ge Zhang, Wenhao Huang, Di He, and Tianle Cai. In-place test-time training. InThe Fourteenth International Conference on Learning Representations, 2026. URL https://openreview.net/forum?id=dTWfCLSoyl. 12
2026
-
[21]
Chelsea Finn, Pieter Abbeel, and Sergey Levine. Model-agnostic meta-learning for fast adaptation of deep networks.arXiv preprint arXiv: 1703.03400, 2017. 5
Pith/arXiv arXiv 2017
-
[22]
In-context imitation learning via next-token prediction.arXiv preprint arXiv:2408.15980, 2024
Letian Fu, Huang Huang, Gaurav Datta, Lawrence Yunliang Chen, William Chung-Ho Panitch, Fangchen Liu, Hui Li, and Ken Goldberg. In-context imitation learning via next-token prediction.arXiv preprint arXiv:2408.15980, 2024. 3, 11
Pith/arXiv arXiv 2024
-
[23]
Gated memory policy.arXiv preprint arXiv: 2604.18933, 2026
Yihuai Gao, Jinyun Liu, Shuang Li, and Shuran Song. Gated memory policy.arXiv preprint arXiv: 2604.18933, 2026. 11, 12
Pith/arXiv arXiv 2026
-
[24]
Vit3: Unlocking test-time training in vision.arXiv preprint arXiv: 2512.01643, 2025
Dongchen Han, Yining Li, Tianyu Li, Zixuan Cao, Ziming Wang, Jun Song, Yu Cheng, Bo Zheng, and Gao Huang. Vit3: Unlocking test-time training in vision.arXiv preprint arXiv: 2512.01643, 2025. 12
Pith/arXiv arXiv 2025
-
[25]
Gaussian error linear units (gelus).arXiv preprint arXiv: 1606.08415,
Dan Hendrycks and Kevin Gimpel. Gaussian error linear units (gelus).arXiv preprint arXiv: 1606.08415,
-
[26]
Long short-term memory.Neural Comput., 9(8):1735–1780, November 1997
Sepp Hochreiter and Jürgen Schmidhuber. Long short-term memory.Neural Comput., 9(8):1735–1780, November 1997. ISSN 0899-7667. doi: 10.1162/neco.1997.9.8.1735. URLhttps://doi.org/10.1162/ neco.1997.9.8.1735. 11
-
[27]
Cheng-Ping Hsieh, Simeng Sun, Samuel Kriman, Shantanu Acharya, Dima Rekesh, Fei Jia, Yang Zhang, and Boris Ginsburg. Ruler: What’s the real context size of your long-context language models?arXiv preprint arXiv: 2404.06654, 2024. 2
Pith/arXiv arXiv 2024
-
[28]
Shengding Hu, Yuge Tu, Xu Han, Chaoqun He, Ganqu Cui, Xiang Long, Zhi Zheng, Yewei Fang, Yuxiang Huang,WeilinZhao,XinrongZhang,ZhengLengThai,KaihuoZhang,ChongyiWang,YuanYao,Chenyang Zhao, Jie Zhou, Jie Cai, Zhongwu Zhai, Ning Ding, Chao Jia, Guoyang Zeng, Dahai Li, Zhiyuan Liu, and Maosong Sun. Minicpm: Unveiling the potential of small language models wit...
Pith/arXiv arXiv 2024
-
[29]
Physical Intelligence, Kevin Black, Noah Brown, James Darpinian, Karan Dhabalia, Danny Driess, Adnan Esmail, Michael Equi, Chelsea Finn, Niccolo Fusai, Manuel Y. Galliker, Dibya Ghosh, Lachy Groom, Karol Hausman, Brian Ichter, Szymon Jakubczak, Tim Jones, Liyiming Ke, Devin LeBlanc, Sergey Levine, Adrian Li-Bell, Mohith Mothukuri, Suraj Nair, Karl Pertsch...
Pith/arXiv arXiv 2025
-
[30]
Perceiver: General perception with iterative attention.arXiv preprint arXiv: 2103.03206, 2021
Andrew Jaegle, Felix Gimeno, Andrew Brock, Andrew Zisserman, Oriol Vinyals, and Joao Carreira. Perceiver: General perception with iterative attention.arXiv preprint arXiv: 2103.03206, 2021. 4
Pith/arXiv arXiv 2021
-
[31]
Huiwon Jang, Sihyun Yu, Heeseung Kwon, Hojin Jeon, Younggyo Seo, and Jinwoo Shin. Contextvla: Vision-language-action model with amortized multi-frame context.arXiv preprint arXiv: 2510.04246,
-
[32]
Vima: General robot manipulation with multimodal prompts.arXiv preprint arXiv: 2210.03094, 2022
Yunfan Jiang, Agrim Gupta, Zichen Zhang, Guanzhi Wang, Yongqiang Dou, Yanjun Chen, Li Fei-Fei, Anima Anandkumar, Yuke Zhu, and Linxi Fan. Vima: General robot manipulation with multimodal prompts.arXiv preprint arXiv: 2210.03094, 2022. 3, 11, 12
Pith/arXiv arXiv 2022
-
[33]
Muon: An optimizer for hidden layers in neural networks, 2024
Keller Jordan, Yuchen Jin, Vlado Boza, Jiacheng You, Franz Cesista, Laker Newhouse, and Jeremy Bernstein. Muon: An optimizer for hidden layers in neural networks, 2024. URLhttps://kellerjordan. github.io/posts/muon/. 20
2024
-
[34]
Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei
Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. Scaling laws for neural language models.arXiv preprint arXiv: 2001.08361, 2020. 11
Pith/arXiv arXiv 2001
-
[35]
Lattice: Learning to efficiently compress the memory.arXiv preprint arXiv: 2504.05646, 2025
Mahdi Karami, Razvan Pascanu, and Vahab Mirrokni. Lattice: Learning to efficiently compress the memory.arXiv preprint arXiv: 2504.05646, 2025. 12
arXiv 2025
-
[36]
OpenVLA: An open-source vision-language- action model
Moo Jin Kim, Karl Pertsch, Siddharth Karamcheti, Ted Xiao, Ashwin Balakrishna, Suraj Nair, Rafael Rafailov, Ethan P Foster, Pannag R Sanketi, Quan Vuong, Thomas Kollar, Benjamin Burchfiel, Russ Tedrake, Dorsa Sadigh, Sergey Levine, Percy Liang, and Chelsea Finn. OpenVLA: An open-source vision-language- action model. In8th Annual Conference on Robot Learni...
2024
-
[37]
Moo Jin Kim, Yihuai Gao, Tsung-Yi Lin, Yen-Chen Lin, Yunhao Ge, Grace Lam, Percy Liang, Shuran Song, Ming-Yu Liu, Chelsea Finn, and Jinwei Gu. Cosmos policy: Fine-tuning video models for visuomotor control and planning.arXiv preprint arXiv: 2601.16163, 2026. 2, 11
Pith/arXiv arXiv 2026
-
[38]
Rma: Rapid motor adaptation for legged robots.arXiv preprint arXiv:2107.04034, 2021
Ashish Kumar, Zipeng Fu, Deepak Pathak, and Jitendra Malik. Rma: Rapid motor adaptation for legged robots.arXiv preprint arXiv:2107.04034, 2021. 11
Pith/arXiv arXiv 2021
-
[39]
In-context reinforcement learning with algorithm distillation
Michael Laskin, Luyu Wang, Junhyuk Oh, Emilio Parisotto, Stephen Spencer, Richie Steigerwald, DJ Strouse, Steven Stenberg Hansen, Angelos Filos, Ethan Brooks, Maxime Gazeau, Himanshu Sahni, Satinder Singh, and Volodymyr Mnih. In-context reinforcement learning with algorithm distillation. In The Eleventh International Conference on Learning Representations...
2023
-
[40]
Causal world modeling for robot control.arXiv preprint arXiv: 2601.21998, 2026
Lin Li, Qihang Zhang, Yiming Luo, Shuai Yang, Ruilin Wang, Fei Han, Mingrui Yu, Zelin Gao, Nan Xue, Xing Zhu, Yujun Shen, and Yinghao Xu. Causal world modeling for robot control.arXiv preprint arXiv: 2601.21998, 2026. 11
Pith/arXiv arXiv 2026
-
[41]
Unified video action model.arXiv preprint arXiv: 2503.00200, 2025
Shuang Li, Yihuai Gao, Dorsa Sadigh, and Shuran Song. Unified video action model.arXiv preprint arXiv: 2503.00200, 2025. 2, 11
Pith/arXiv arXiv 2025
-
[42]
Tnt: Improving chunkwise training for test-time memorization.arXiv preprint arXiv:2511.07343, 2025
Zeman Li, Ali Behrouz, Yuan Deng, Peilin Zhong, Praneeth Kacham, Mahdi Karami, Meisam Razaviyayn, and Vahab Mirrokni. Tnt: Improving chunkwise training for test-time memorization.arXiv preprint arXiv:2511.07343, 2025. 12
arXiv 2025
-
[43]
Parallelizing non-linear sequential models over the sequence length
Yi Heng Lim, Qi Zhu, Joshua Selfridge, and Muhammad Firmansyah Kasim. Parallelizing non-linear sequential models over the sequence length. InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. URLhttps: //openreview.net/forum?id=E34AlVLN0v. 12
2024
-
[44]
Fanqi Lin, Ruiqian Nai, Yingdong Hu, Jiacheng You, Junming Zhao, and Yang Gao. Onetwovla: A unified vision-language-action model with adaptive reasoning.arXiv preprint arXiv:2505.11917, 2025. 11
arXiv 2025
-
[45]
On-the-fly vla adaptation via test-time reinforcement learning.arXiv preprint arXiv:2601.06748, 2026
Changyu Liu, Yiyang Liu, Taowen Wang, Qiao Zhuang, James Chenhao Liang, Wenhao Yang, Renjing Xu, Qifan Wang, Dongfang Liu, and Cheng Han. On-the-fly vla adaptation via test-time reinforcement learning.arXiv preprint arXiv:2601.06748, 2026. 12 16 RoboTTT: Context Scaling for Robot Policies
Pith/arXiv arXiv 2026
-
[46]
Lost in the middle: How language models use long contexts.Transactions of the association for computational linguistics, 12:157–173, 2024
Nelson F Liu, Kevin Lin, John Hewitt, Ashwin Paranjape, Michele Bevilacqua, Fabio Petroni, and Percy Liang. Lost in the middle: How language models use long contexts.Transactions of the association for computational linguistics, 12:157–173, 2024. 2
2024
-
[47]
RDT-1B:a diffusionfoundationmodelforbimanual manipulation
Songming Liu, Lingxuan Wu, Bangguo Li, Hengkai Tan, Huayu Chen, Zhengyi Wang, Ke Xu, Hang Su, and JunZhu. RDT-1B:a diffusionfoundationmodelforbimanual manipulation. InTheThirteenthInternational Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025. URLhttps://openreview.net/forum?id=yAzN4tz7oI. 2, 11
2025
-
[48]
Loshchilov and F
I. Loshchilov and F. Hutter. Decoupled weight decay regularization.International Conference on Learning Representations, 2017. 20
2017
-
[49]
Savarese, Yuke Zhu, and Roberto Mart’in-Mart’in
Ajay Mandlekar, Danfei Xu, Josiah Wong, Soroush Nasiriany, Chen Wang, Rohun Kulkarni, Li Fei-Fei, S. Savarese, Yuke Zhu, and Roberto Mart’in-Mart’in. What matters in learning from offline human demonstrations for robot manipulation.Conference on Robot Learning, 2021. 11
2021
-
[50]
Max Sobol Mark, Jacky Liang, Maria Attarian, Chuyuan Fu, Debidatta Dwibedi, Dhruv Shah, and Aviral Kumar. Bpp: Long-context robot imitation learning by focusing on key history frames.arXiv preprint arXiv: 2602.15010, 2026. 2, 11, 12
arXiv 2026
-
[51]
NVIDIA, Johan Bjorck, Fernando Castañeda, Nikita Cherniadev, Xingye Da, Runyu Ding, Linxi "Jim" Fan, Yu Fang, Dieter Fox, Fengyuan Hu, Spencer Huang, Joel Jang, Zhenyu Jiang, Jan Kautz, Kaushil Kundalia, Lawrence Lao, Zhiqi Li, Zongyu Lin, Kevin Lin, Guilin Liu, Edith Llontop, Loic Magne, Ajay Mandlekar, Avnish Narayan, Soroush Nasiriany, Scott Reed, You ...
Pith/arXiv arXiv 2025
-
[52]
Scalable diffusion models with transformers.arXiv preprint arXiv: 2212.09748, 2022
William Peebles and Saining Xie. Scalable diffusion models with transformers.arXiv preprint arXiv: 2212.09748, 2022. 4, 20
Pith/arXiv arXiv 2022
-
[53]
Yarn: Efficient context window extension of large language models
Bowen Peng, Jeffrey Quesnelle, Honglu Fan, and Enrico Shippole. Yarn: Efficient context window extension of large language models. InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. URLhttps://openreview.net/ forum?id=wHBfxhZu1u. 2
2024
-
[54]
Fast: Efficient action tokenization for vision-language-action models, 2025
Karl Pertsch, Kyle Stachowicz, Brian Ichter, Danny Driess, Suraj Nair, Quan Vuong, Oier Mees, Chelsea Finn, and Sergey Levine. Fast: Efficient action tokenization for vision-language-action models, 2025. URL https://arxiv.org/abs/2501.09747. 12
Pith/arXiv arXiv 2025
-
[55]
In-hand object rotation via rapid motor adaptation
Haozhi Qi, Ashish Kumar, Roberto Calandra, Yi Ma, and Jitendra Malik. In-hand object rotation via rapid motor adaptation. InConference on Robot Learning, pages 1722–1732. PMLR, 2023. 11
2023
-
[56]
Scott E. Reed, Konrad Zolna, Emilio Parisotto, Sergio Gómez Colmenarejo, Alexander Novikov, Gabriel Barth-Maron, Mai Gimenez, Yury Sulsky, Jackie Kay, Jost Tobias Springenberg, Tom Eccles, Jake Bruce, Ali Razavi, Ashley Edwards, Nicolas Heess, Yutian Chen, Raia Hadsell, Oriol Vinyals, Mahyar Bordbar, and Nando de Freitas. A generalist agent.Trans. Mach. L...
2022
-
[57]
A reduction of imitation learning and structured prediction to no-regret online learning
Stephane Ross, Geoffrey Gordon, and Drew Bagnell. A reduction of imitation learning and structured prediction to no-regret online learning. In Geoffrey Gordon, David Dunson, and Miroslav Dudík, editors, Proceedings of the Fourteenth International Conference on Artificial Intelligence and Statistics, volume 15 of Proceedings of Machine Learning Research, p...
2011
-
[58]
Linear transformers are secretly fast weight programmers
Imanol Schlag, Kazuki Irie, and Jürgen Schmidhuber. Linear transformers are secretly fast weight programmers. InInternational conference on machine learning, pages 9355–9366. PMLR, 2021. 2 17 RoboTTT: Context Scaling for Robot Policies
2021
-
[59]
Eagle: Exploring the design space for multimodal llms with mixture of encoders
Min Shi, Fuxiao Liu, Shihao Wang, Shijia Liao, Subhashree Radhakrishnan, De-An Huang, Hongxu Yin, Karan Sapra, Yaser Yacoob, Humphrey Shi, Bryan Catanzaro, Andrew Tao, Jan Kautz, Zhiding Yu, and Guilin Liu. Eagle: Exploring the design space for multimodal llms with mixture of encoders. InICLR,
-
[60]
Smolvla: A vision-language-action model for affordable and efficient robotics
Mustafa Shukor, Dana Aubakirova, Francesco Capuano, Pepijn Kooijmans, Steven Palma, Adil Zouitine, Michel Aractingi, Caroline Pascal, Martino Russi, Andres Marafioti, Simon Alibert, Matthieu Cord, Thomas Wolf, and Remi Cadene. Smolvla: A vision-language-action model for affordable and efficient robotics. arXiv preprint arXiv: 2506.01844, 2025. 2, 11
Pith/arXiv arXiv 2025
-
[61]
Ajay Sridhar, Jennifer Pan, Satvik Sharma, and Chelsea Finn. Memer: Scaling up memory for robot control via experience retrieval.arXiv preprint arXiv: 2510.20328, 2025. 11, 12
arXiv 2025
-
[62]
Roformer: Enhanced transformer with rotary position embedding.arXiv preprint arXiv: 2104.09864, 2021
Jianlin Su, Yu Lu, Shengfeng Pan, Ahmed Murtadha, Bo Wen, and Yunfeng Liu. Roformer: Enhanced transformer with rotary position embedding.arXiv preprint arXiv: 2104.09864, 2021. 20
Pith/arXiv arXiv 2021
-
[63]
Efros, and Moritz Hardt
Yu Sun, Xiaolong Wang, Zhuang Liu, John Miller, Alexei A. Efros, and Moritz Hardt. Test-time training with self-supervision for generalization under distribution shifts. InProceedings of the 37th International Conference on Machine Learning, ICML’20. JMLR.org, 2020. 12, 20
2020
-
[64]
Yu Sun, Xinhao Li, Karan Dalal, Jiarui Xu, Arjun Vikram, Genghan Zhang, Yann Dubois, Xinlei Chen, Xiaolong Wang, Sanmi Koyejo, Tatsunori Hashimoto, and Carlos Guestrin. Learning to (learn at test time): Rnns with expressive hidden states.arXiv preprint arXiv: 2407.04620, 2024. 2, 3
Pith/arXiv arXiv 2024
-
[65]
End-to-end test-time training for long context.arXiv preprint arXiv: 2512.23675, 2025
Arnuv Tandon, Karan Dalal, Xinhao Li, Daniel Koceja, Marcel Rød, Sam Buchanan, Xiaolong Wang, Jure Leskovec, Sanmi Koyejo, Tatsunori Hashimoto, Carlos Guestrin, Jed McCaleb, Yejin Choi, and Yu Sun. End-to-end test-time training for long context.arXiv preprint arXiv: 2512.23675, 2025. 5, 12
arXiv 2025
-
[66]
Human-timescale adaptation in an open-ended task space.arXiv preprint arXiv:2301.07608, 2023
Adaptive Agent Team, Jakob Bauer, Kate Baumli, Satinder Baveja, Feryal Behbahani, Avishkar Bhoopc- hand, Nathalie Bradley-Schmieg, Michael Chang, Natalie Clay, Adrian Collister, et al. Human-timescale adaptation in an open-ended task space.arXiv preprint arXiv:2301.07608, 2023. 11
Pith/arXiv arXiv 2023
-
[67]
Sanketi, Quan Vuong, Ted Xiao, Dorsa Sadigh, Chelsea Finn, and Sergey Levine
Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Tobias Kreiman, Charles Xu, Jianlan Luo, You Liang Tan, Pannag R. Sanketi, Quan Vuong, Ted Xiao, Dorsa Sadigh, Chelsea Finn, and Sergey Levine. Octo: An open-source generalist robot policy. ROBOTICS, 2024. doi: 10.48550/arXiv.2405.12213. URLhttps://...
-
[68]
Marcel Torne, Andy Tang, Yuejiang Liu, and Chelsea Finn. Learning long-context diffusion policies via past-token prediction.arXiv preprint arXiv: 2505.09561, 2025. 11, 12
Pith/arXiv arXiv 2025
-
[69]
Marcel Torne, Karl Pertsch, Homer Walke, Kyle Vedder, Suraj Nair, Brian Ichter, Allen Z. Ren, Haohuan Wang, Jiaming Tang, Kyle Stachowicz, Karan Dhabalia, Michael Equi, Quan Vuong, Jost Tobias Springen- berg, Sergey Levine, Chelsea Finn, and Danny Driess. Mem: Multi-scale embodied memory for vision language action models.arXiv preprint arXiv: 2603.03596, ...
arXiv 2026
-
[70]
Chuan Wen, Jierui Lin, Trevor Darrell, Dinesh Jayaraman, and Yang Gao. Fighting copycat agents in behavioral cloning from observation histories.arXiv preprint arXiv: 2010.14876, 2020. 11
Pith/arXiv arXiv 2010
-
[71]
Efficient streaming language models with attention sinks
Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. Efficient streaming language models with attention sinks. InThe Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024. URLhttps://openreview.net/forum? id=NG7sS51zVF. 2
2024
-
[72]
Magma: A foundation model for multimodal ai agents, 2025
Jianwei Yang, Reuben Tan, Qianhui Wu, Ruijie Zheng, Baolin Peng, Yongyuan Liang, Yu Gu, Mu Cai, Seonghyeon Ye, Joel Jang, Yuquan Deng, Lars Liden, and Jianfeng Gao. Magma: A foundation model for multimodal ai agents, 2025. URLhttps://arxiv.org/abs/2502.13130. 12 18 RoboTTT: Context Scaling for Robot Policies
Pith/arXiv arXiv 2025
-
[73]
Fla: A triton-based library for hardware-efficient implementations of linear attention mechanism, January 2024
Songlin Yang and Yu Zhang. Fla: A triton-based library for hardware-efficient implementations of linear attention mechanism, January 2024. URLhttps://github.com/fla-org/flash-linear-attention. 22
2024
-
[74]
Songlin Yang, Jan Kautz, and Ali Hatamizadeh. Gated delta networks: Improving mamba2 with delta rule.International Conference on Learning Representations, 2024. doi: 10.48550/arXiv.2412.06464. 7
-
[75]
World action models are zero-shot policies.arXiv preprint arXiv:2602.15922, 2026
Seonghyeon Ye, Yunhao Ge, Kaiyuan Zheng, Shenyuan Gao, Sihyun Yu, George Kurian, Suneel Indupuru, You Liang Tan, Chuning Zhu, Jiannan Xiang, et al. World action models are zero-shot policies.arXiv preprint arXiv:2602.15922, 2026. 2, 11
Pith/arXiv arXiv 2026
-
[76]
Tianyuan Yuan, Zibin Dong, Yicheng Liu, and Hang Zhao. Fast-wam: Do world action models need test-time future imagination?arXiv preprint arXiv: 2603.16666, 2026. 2, 11
Pith/arXiv arXiv 2026
-
[77]
Learning to discover at test time, 2026
Mert Yuksekgonul, Daniel Koceja, Xinhao Li, Federico Bianchi, Jed McCaleb, Xiaolong Wang, Jan Kautz, Yejin Choi, James Zou, Carlos Guestrin, and Yu Sun. Learning to discover at test time, 2026. URL https://arxiv.org/abs/2601.16175. 12
Pith/arXiv arXiv 2026
-
[78]
Tianyuan Zhang, Sai Bi, Yicong Hong, Kai Zhang, Fujun Luan, Songlin Yang, Kalyan Sunkavalli, William T. Freeman, and Hao Tan. Test-time training done right.arXiv preprint arXiv: 2505.23884, 2025. 2, 3, 11, 12, 20
Pith/arXiv arXiv 2025
-
[79]
Fast-weight product key memory.arXiv preprint arXiv: 2601.00671, 2026
Tianyu Zhao and Llion Jones. Fast-weight product key memory.arXiv preprint arXiv: 2601.00671, 2026. 12
arXiv 2026
-
[80]
TraceVLA: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies
Ruijie Zheng, Yongyuan Liang, Shuaiyi Huang, Jianfeng Gao, Hal Daumé III, Andrey Kolobov, Furong Huang, and Jianwei Yang. TraceVLA: Visual trace prompting enhances spatial-temporal awareness for generalist robotic policies. InThe Thirteenth International Conference on Learning Representations, 2025. 11, 12
2025
discussion (0)
Sign in with ORCID, Apple, or X to comment. Anyone can read and Pith papers without signing in.