REVIEW 4 major objections 4 minor 1 cited by
VertiFormer: A Data-Efficient Multi-Task Transformer for Off-Road Robot Mobility
T0 review · 4 major / 4 minor · reviewed 2026-08-09 · deepseek-v4-flash
Pith's one-line read VertiFormer predicts next pose, action, and terrain patch from one hour of driving data, and beats a terrain-specialist kinodynamic model on a physical robot.
desk verdict A genuinely new multi-task transformer for off-road kinodynamics, with a load-bearing data-claim contradiction that must be fixed before the efficiency story holds. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The central object is the unified multi-modal latent token $z_t = f_s(\hat a_t \cdot \hat p_t \cdot \hat i_t)$, a learned linear projection that concatenates embedded actions, poses, and terrain patches into one homogeneous token stream before the encoder. This shared token space is the inductive bias that the paper credits for letting one hour of data carry three tasks. Training alternates two learned masks with equal probability: action-conditioned pose prediction for forward kinodynamics, and pose-conditioned action prediction for inverse kinodynamics, while masking both enables zero-shot behavior cloning. Multi-context tokens and cross-attention with causal masking let the decoder emit all future steps non-autoregressively, avoiding the error accumulation that autoregressive decoding suffers over long horizons.
What would settle it
Train the same multi-task masking and non-autoregressive decoder with separate per-modality tokens instead of the unified latent projection, and compare held-out forward, inverse, and behavior-cloning error on the same one-hour dataset. If error rates are statistically indistinguishable, the unified representation is not the cause of the data-efficiency gain. Also, count how many rollover and high-roll episodes occur in the one-hour log; if those states appear only a handful of times, the claim that the log covers dangerous kinodynamic interactions is weak.
Extended reading notes
Core claim
On the paper's own terms, the discovery is that a Transformer can be made data-efficient for off-road mobility by replacing separate per-modality tokens with a unified latent representation and by alternating learnable masks between future actions and future poses during training. Trained on one hour of human teleoperation over a rock testbed, VertiFormer predicts the next pose, action, and terrain patch simultaneously, and its forward kinodynamic model, when plugged into a sampling-based planner, completes ten of ten traversals on an unseen terrain testbed while a state-of-the-art kinodynamic model specialized for vertically challenging terrain completes eight of ten. The same architecture also produces inverse kinodynamics and zero-shot behavior cloning without retraining dedicated heads.
Load-bearing premise
The load-bearing premise is that combining pose, action, and terrain into one latent token makes all three tasks share a representation that one hour of data can learn; the paper states this as a hypothesis and supports it mainly with a temporal-order classification ablation rather than direct downstream-task evidence.
Editorial extensions
If this is right
- A single one-hour teleoperation log can be enough to train a forward kinodynamic model that sampling-based planners can use to traverse rocky, vertically challenging terrain.
- The same trained weights handle inverse kinodynamics and zero-shot behavior cloning, so a robot can switch tasks or tolerate missing future pose or action inputs without retraining separate heads.
- Because the decoder is non-autoregressive, one-second and two-second predictions drift less than an autoregressive decoder's, which matters for real-time planning loops.
- Under extreme data scarcity, sinusoidal positional encoding and a final normalization layer help, while an auxiliary terrain-patch reconstruction head can degrade the primary tasks.
Reading between the lines
- I infer that the equal 50/50 split between action-conditioned and pose-conditioned masking is a design choice, not a tuned optimum; varying this ratio and the mask token's learning rate could change how quickly shared representations emerge from small data.
- I infer that the unified-token projection could transfer to other locomotion forms, such as legged robots or tracked vehicles, whenever states and commands admit vector embeddings; the paper only demonstrates wheeled motion.
- I infer that the temporal-order classification result is a proxy, not a measure of downstream task quality; a cleaner test would ablate unified versus separate tokens and report forward, inverse, and cloning error on held-out terrain.
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. VertiFormer is a Transformer architecture for off-road mobility that combines a Transformer encoder and decoder, a unified latent representation for actions, poses, and terrain patches, learnable masking, and non-autoregressive multi-token prediction. It is trained on human teleoperation data claimed to be one hour long and evaluated on forward and inverse kinodynamic modeling and behavior cloning, first as offline prediction error against TAL and VertiEncoder, then on a physical V4W robot with MPPI, Dijkstra, and behavior-cloning controllers. The paper reports ablations on positional encoding, normalization, unified representation, patch head, prediction horizon, and training paradigm. The stated contributions are a data-efficient multi-task Transformer, empirical design guidelines, and physical-robot demonstrations.
Significance. If the central claims held, this would be a practically relevant demonstration that a Transformer can learn multiple off-road mobility tasks from roughly one hour of demonstration data and run onboard a robot, with a released open-source implementation. The strengths of the paper are the real-robot evaluation, the systematic ablation structure, and the clear architectural description. However, the headline data-efficiency claim is currently undercut by the dataset description in Appendix B, and the numerical comparisons lack error bars, confidence intervals, or significance tests. The evidence for the central multi-task mechanism is indirect. These issues are fixable but currently prevent the paper's claims from being fully verified.
major comments (4)
- [Appendix B vs. Abstract/§IV] The paper repeatedly states that VertiFormer is 'trained with only one hour of data' (Abstract, §I, §IV, §IV-C), but Appendix B, the only dataset description, says: 'The dataset includes 30 minutes of data from both a planar surface and the rock testbed.' This is ambiguous between 30 minutes total and 30 minutes per surface, but in either reading it is not a statement of one hour. Because the data-efficiency claim and the comparison with TAL [19] depend on the exact training-set size and composition, the authors must reconcile this discrepancy, report the exact duration, the proportion of planar vs. rock-teleoperation data, and the precise train/test split, and state whether TAL and all other baselines were trained on the identical data. Without this audit, the headline result cannot be evaluated.
- [§III-B and Table I] The headline offline comparison reports error rates of 0.495 (VertiFormer), 0.516 (Nazeri et al.), and 0.528 (TAL) with no confidence intervals, number of test trajectories, or significance tests; the 0.033 difference between VertiFormer and TAL is small and could be sampling noise. Similarly, Table I reports success rates out of 10 physical trials (e.g., FKD: TAL 8/10 vs. VertiFormer 10/10) without confidence intervals, per-trial distributions, or a statistical test, and the mean roll/pitch columns include wide standard deviations. The claim that VertiFormer 'outperforms' TAL on navigation performance is therefore not statistically established. Please report repeated-seed or bootstrap intervals for the offline metrics, per-trial results for Table I, and an appropriate significance test or an explicit statement that the differences are descriptive only.
- [§IV-B / Fig. 5 and §III-A2] The only direct evidence offered for the unified-latent-representation benefit is Fig. 5, an auxiliary temporal-order classification loss, shown as a single learning curve with no error bars and no downstream FKD/IKD/BC results. The paper itself labels the multi-task data-efficiency mechanism as 'hypothesized' in §III-A2. To support the central claim that alternating masking plus unified representation drives the data-efficiency gain, the authors should show the unified vs. separate-modality comparison on the actual downstream tasks and include variance over training seeds. As it stands, the paper does not establish that the gains come from multi-task sharing rather than from the unified input representation alone.
- [§IV-C / Fig. 7 and Abstract] The abstract and introduction state that VertiFormer predicts the next pose, action, and terrain patch, but Fig. 7 shows that adding the patch reconstruction head degrades performance on the primary tasks, and the text explains that the patch head 'introduces noise into the learning process.' It is never stated whether the final deployed VertiFormer includes the patch head. If it does not, the terrain-patch prediction contribution is not part of the final model; if it does, the ablation appears to contradict the design. The authors should state the final architecture explicitly, including which heads are used in the robot experiments, and align the abstract and introduction with that choice.
minor comments (4)
- [Throughout] The notation 'V ERTI FORMER', 'V ERTI ENCODER', and 'V ERTI DECODER' is typeset with irregular spacing; use a consistent single-token name throughout.
- [Appendix A, Table III] Table III cites [73] (Sennrich et al., subword units) for batch normalization; this appears to be the wrong reference and should be corrected.
- [§III-B] The offline error-rate table has no table number or caption; add one and specify how the error rate is computed (e.g., normalized by what quantity) and over how many held-out samples.
- [Fig. 6] Fig. 6 includes a panel labeled 'IKD' alongside X/Y/Z panels, but the IKD metric is not defined in the caption; clarify what is plotted in each panel and the units.
Circularity Check
No significant circularity: the reported predictions are evaluated on held-out data, and the self-citations are empirical comparators rather than load-bearing derivation inputs.
full rationale
The paper's derivation chain is self-contained with respect to circularity. VertiFormer's FKD, IKD, and BC outputs are obtained by training the model on a fixed dataset and measuring prediction error against ground truth, with no parameter fitted directly to the reported success metrics and then renamed as a prediction. The unified latent representation, learnable masking, and non-autoregressive training are described as architectural design choices, and the multi-task benefit is explicitly labeled a hypothesis in Section III-A2 rather than being assumed through an equation that defines the output in terms of the input. The baselines TAL [19] and VertiEncoder [56] are prior works with overlapping authorship, but they are used as experimental comparators for physical-robot outcomes, not as justifications for why VertiFormer's architecture must work; no uniqueness theorem or self-citation chain forces the model choice. The paper does contain a notable internal inconsistency: the main text repeatedly claims 'one hour' of training data, while Appendix B states 'The dataset includes 30 minutes of data from both a planar surface and the rock testbed.' This is a reporting and audit concern about the size of the training set, and it may undermine the quantitative force of the data-efficiency claim, but it is not a circular reduction of a prediction to its inputs. No equation in the paper equates a claimed result with a fitted parameter or defines an output in terms of the target quantity.
Assumptions & free parameters
free parameters (7)
- hidden size D =
512
- encoder/decoder layers =
6 encoder, 4 decoder
- attention heads =
8
- dropout =
0.3
- learning rate and weight decay =
5e-4, 0.08
- batch size and epochs =
512, 200
- prediction horizon tau =
3 steps at 3 Hz
assumptions (4)
- domain assumption Transformer self-attention can learn kinodynamic mappings from sequential robot data when inputs are projected into a unified embedding.
- domain assumption Elevation-map terrain patches and 6-DoF odometry contain sufficient information to model tire deformation, suspension travel, and wheel-terrain interaction.
- domain assumption Human teleoperation on the rock testbed is a valid expert policy for training FKD, IKD, and BC.
- ad hoc to paper Alternating action and pose masking improves data efficiency by forcing shared multi-task latent representations.
Cite this review
Pith. "Pith review of VertiFormer: A Data-Efficient Multi-Task Transformer for Off-Road Robot Mobility." pith.science (2026). https://pith.science/paper/QDJ7KOVT
@misc{pith2026250200543,
author = {Pith},
title = {Pith review of: VertiFormer: A Data-Efficient Multi-Task Transformer for Off-Road Robot Mobility},
year = {2026},
howpublished = {\url{https://pith.science/paper/QDJ7KOVT}},
note = {Machine review of arXiv:2502.00543}
}
read the original abstract
Sophisticated learning architectures, e.g., Transformers, present a unique opportunity for robots to understand complex vehicle-terrain kinodynamic interactions for off-road mobility. While internet-scale data are available for Natural Language Processing (NLP) and Computer Vision (CV) tasks to train Transformers, real-world mobility data are difficult to acquire with physical robots navigating off-road terrain. Furthermore, training techniques specifically designed to process text and image data in NLP and CV may not apply to robot mobility. In this paper, we propose VertiFormer, a novel data-efficient multi-task Transformer model trained with only one hour of data to address such challenges of applying Transformer architectures for robot mobility on extremely rugged, vertically challenging, off-road terrain. Specifically, VertiFormer employs a new learnable masked modeling and next token prediction paradigm to predict the next pose, action, and terrain patch to enable a variety of off-road mobility tasks simultaneously, e.g., forward and inverse kinodynamics modeling. The non-autoregressive design mitigates computational bottlenecks and error propagation associated with autoregressive models. VertiFormer's unified modality representation also enhances learning of diverse temporal mappings and state representations, which, combined with multiple objective functions, further improves model generalization. Our experiments offer insights into effectively utilizing Transformers for off-road robot mobility with limited data and demonstrate our efficiently trained Transformer can facilitate multiple off-road mobility tasks onboard a physical mobile robot.
Figures
Figures from the paper (9 more)
Forward citations
Cited by 1 Pith paper
-
Narrate2Nav: Real-Time Visual Navigation with Implicit Language Reasoning in Human-Centric Environments
Narrate2Nav uses Barlow Twins alignment to distill language-based reasoning from a large teacher into a small RGB-only navigation model, reporting lower trajectory error and higher goal-reaching success than four baselines.
Reference graph
Works this paper leans on
-
[19]
Aniket Datar, Chenhui Pan, Mohammad Nazeri, Anuj Pokhrel, and Xuesu Xiao. Terrain-Attentive Learning for Efficient 6-DoF Kinodynamic Modeling on Vertically Challenging Terrain. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages 5438–5443, Abu Dhabi, United Arab Emirates, October 2024. IEEE. ISBN 979-8-3503-7770-5. d...
-
[1]
Invariance is Key to Generalization: Examining the Role of Representation in Sim-to-Real Transfer for Visual Navigation, December 2023
Bo Ai, Zhanxin Wu, and David Hsu. Invariance is Key to Generalization: Examining the Role of Representation in Sim-to-Real Transfer for Visual Navigation, December 2023
2023
-
[2]
Self-supervised learning from images with a joint-embedding predictive architecture
Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bo- janowski, Pascal Vincent, Michael Rabbat, Yann LeCun, and Nicolas Ballas. Self-supervised learning from images with a joint-embedding predictive architecture. arXiv preprint arXiv:2301.08243, 2023
arXiv 2023
-
[3]
Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer Normalization, July 2016
2016
-
[4]
Curriculum learning for vehicle lateral stability estimations
Jihwan Bae, Taekyung Kim, Wonsuk Lee, and Inwook Shim. Curriculum learning for vehicle lateral stability estimations. IEEE Access, 9:89249–89262, 2021
2021
-
[5]
Navigation World Models, December 2024
Amir Bar, Gaoyue Zhou, Danny Tran, Trevor Darrell, and Yann LeCun. Navigation World Models, December 2024
2024
-
[6]
MC- JEPA: A Joint-Embedding Predictive Architecture for Self-Supervised Learning of Motion and Content Fea- tures, July 2023
Adrien Bardes, Jean Ponce, and Yann LeCun. MC- JEPA: A Joint-Embedding Predictive Architecture for Self-Supervised Learning of Motion and Content Fea- tures, July 2023
2023
-
[7]
Revisiting Feature Prediction for Learning Visual Representations from Video, February 2024
Adrien Bardes, Quentin Garrido, Jean Ponce, Xinlei Chen, Michael Rabbat, Yann LeCun, Mahmoud Assran, and Nicolas Ballas. Revisiting Feature Prediction for Learning Visual Representations from Video, February 2024
2024
Show all 97 references
-
[8]
Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, Xin Zhang, Jake Zhao, and Karol Zieba
Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D. Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, Xin Zhang, Jake Zhao, and Karol Zieba. End to End Learning for Self-Driving Cars, April 2016
2016
-
[9]
A survey on terrain traversability analysis for autonomous ground ve- hicles: Methods, sensors, and challenges
Paulo Borges, Thierry Peynot, Sisi Liang, Bilal Arain, Matt Wildie, Melih Minareci, Serge Lichman, Garima Samvedi, Inkyu Sa, Nicolas Hudson, Michael Milford, Peyman Moghadam, and Peter Corke. A survey on terrain traversability analysis for autonomous ground ve- hicles: Methods...
2022 doi
-
[10]
Language models are few-shot learners
Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey ...
1901
-
[11]
Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom
Holger Caesar, Varun Bankiti, Alex H. Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A multimodal dataset for autonomous driving. In CVPR, 2020
2020
-
[12]
Evora: Deep evidential traversability learning for risk-aware off-road autonomy
Xiaoyi Cai, Siddharth Ancha, Lakshay Sharma, Philip R Osteen, Bernadette Bucher, Stephen Phillips, Jiuguang Wang, Michael Everett, Nicholas Roy, and Jonathan P How. Evora: Deep evidential traversability learning for risk-aware off-road autonomy. IEEE Transactions on Robotics, 2024
2024
-
[13]
How does it feel? self-supervised costmap learning for off-road vehicle traversability
Mateo Guaman Castro, Samuel Triest, Wenshan Wang, Jason M Gregory, Felix Sanchez, John G Rogers, and Sebastian Scherer. How does it feel? self-supervised costmap learning for off-road vehicle traversability. In 2023 IEEE International Conference on Robotics and Automation (ICR...
2023
-
[14]
Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey, De- cember 2024
Liang Chen, Zekun Wang, Shuhuai Ren, Lei Li, Haozhe Zhao, Yunshui Li, Zefan Cai, Hongcheng Guo, Lei Zhang, Yizhe Xiong, Yichi Zhang, Ruoyu Wu, Qingxiu Dong, Ge Zhang, Jian Yang, Lingwei Meng, Shujie Hu, Yulong Chen, Junyang Lin, Shuai Bai, Andreas Vlachos, Xu Tan, Minjia Zhang...
2024
-
[15]
Decision Trans- former: Reinforcement Learning via Sequence Modeling
Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Ar- avind Srinivas, and Igor Mordatch. Decision Trans- former: Reinforcement Learning via Sequence Modeling. arXiv:2106.01345 [cs], June 2021
2021 arXiv
-
[16]
Generative pretraining from pixels
Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. Generative pretraining from pixels. In Hal Daum ´e III and Aarti Singh, editors, Proceedings of the 37th International Conference on Machine Learning , volume 119 of Pro- ceedings of ...
-
[17]
An Empirical Study of Training Self-Supervised Vision Transformers
Xinlei Chen, Saining Xie, and Kaiming He. An Empirical Study of Training Self-Supervised Vision Transformers. In 2021 IEEE/CVF International Conference on Com- puter Vision (ICCV) , pages 9620–9629, Montreal, QC, Canada, October 2021. IEEE. ISBN 978-1-6654-2812-5. doi: 10.1109...
2021
-
[18]
Hybrid imitative planning with geometric and predictive costs in off-road environ- ments
Nitish Dashora, Daniel Shin, Dhruv Shah, Henry Leopold, David Fan, Ali Agha-Mohammadi, Nicholas Rhinehart, and Sergey Levine. Hybrid imitative planning with geometric and predictive costs in off-road environ- ments. In 2022 International Conference on Robotics and Automation (...
2022
-
[21]
Learning to model and plan for wheeled mobility on vertically challenging terrain
Aniket Datar, Chenhui Pan, and Xuesu Xiao. Learning to model and plan for wheeled mobility on vertically challenging terrain. IEEE Robotics and Automation Letters, 10(2):1505–1512, 2025
2025
-
[22]
BERT: Pre-training of deep bidi- rectional transformers for language understanding
Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidi- rectional transformers for language understanding. In Jill Burstein, Christy Doran, and Thamar Solorio, edi- tors, Proceedings of the 2019 Conference of the North American Chapte...
2019
-
[23]
A note on two problems in connexion with graphs
Edsger W Dijkstra. A note on two problems in connexion with graphs. Numerische mathematik , 1(1):269–271, 1959
1959
-
[24]
Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation, August 2024
Ria Doshi, Homer Walke, Oier Mees, Sudeep Dasari, and Sergey Levine. Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation, August 2024
2024
-
[25]
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, June 2021
Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An Image is Worth 16x16 Words: Transformers for Image Recognition at...
2021
-
[26]
VTNet: Visual Transformer Network for Object Goal Navigation, May 2021
Heming Du, Xin Yu, and Liang Zheng. VTNet: Visual Transformer Network for Object Goal Navigation, May 2021
2021
-
[27]
Tenenbaum, Leslie Kaelbling, Andy Zeng, and Jonathan Tompson
Yilun Du, Mengjiao Yang, Pete Florence, Fei Xia, Ayzaan Wahid, Brian Ichter, Pierre Sermanet, Tianhe Yu, Pieter Abbeel, Joshua B. Tenenbaum, Leslie Kaelbling, Andy Zeng, and Jonathan Tompson. Video Language Planning, October 2023
2023
-
[28]
Step: Stochastic traversability evaluation and planning for risk- aware off-road navigation
David D Fan, Kyohei Otsu, Yuki Kubo, Anushri Dixit, Joel Burdick, and Ali-Akbar Agha-Mohammadi. Step: Stochastic traversability evaluation and planning for risk- aware off-road navigation. In Robotics: Science and Systems (RSS), 2021
2021
-
[29]
Masked Autoencoders As Spatiotemporal Learners, May 2022
Christoph Feichtenhofer, Haoqi Fan, Yanghao Li, and Kaiming He. Masked Autoencoders As Spatiotemporal Learners, May 2022
2022
-
[30]
Foundation Models in Robotics: Applications, Chal- lenges, and the Future
Roya Firoozi, Johnathan Tucker, Stephen Tian, Anirudha Majumdar, Jiankai Sun, Weiyu Liu, Yuke Zhu, Shuran Song, Ashish Kapoor, Karol Hausman, Brian Ichter, Danny Driess, Jiajun Wu, Cewu Lu, and Mac Schwa- ger. Foundation Models in Robotics: Applications, Chal- lenges, and the ...
2023
-
[31]
How to train vision transformer on small-scale datasets? In 33rd British Machine Vision Conference Proceedings, BMVC 2022, 2022
Hanan Gani, Muzammal Naseer, and Mohammad Yaqub. How to train vision transformer on small-scale datasets? In 33rd British Machine Vision Conference Proceedings, BMVC 2022, 2022
2022
-
[32]
Multimodal Masked Autoencoders Learn Transferable Representations, May 2022
Xinyang Geng, Hao Liu, Lisa Lee, Dale Schuurmans, Sergey Levine, and Pieter Abbeel. Multimodal Masked Autoencoders Learn Transferable Representations, May 2022
2022
-
[33]
Zhang, Shaoqing Ren, and Jian Sun
Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2015. doi: 10.1109/cvpr.2016. 90
2016 doi
-
[34]
Masked autoencoders are scalable vision learners
Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ar, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16000–16009, 2022
2022
-
[35]
Bridging nonlineari- ties and stochastic regularizers with gaussian error linear units, 2017
Dan Hendrycks and Kevin Gimpel. Bridging nonlineari- ties and stochastic regularizers with gaussian error linear units, 2017
2017
-
[36]
GAIA-1: A Generative World Model for Autonomous Driving, September 2023
Anthony Hu, Lloyd Russell, Hudson Yeo, Zak Murez, George Fedoseev, Alex Kendall, Jamie Shotton, and Gianluca Corrado. GAIA-1: A Generative World Model for Autonomous Driving, September 2023
2023
-
[37]
DrivingWorld: Constructing World Model for Au- tonomous Driving via Video GPT, December 2024
Xiaotao Hu, Wei Yin, Mingkai Jia, Junyuan Deng, Xi- aoyang Guo, Qian Zhang, Xiaoxiao Long, and Ping Tan. DrivingWorld: Constructing World Model for Au- tonomous Driving via Video GPT, December 2024
2024
-
[38]
Video Prediction Pol- icy: A Generalist Robot Policy with Predictive Visual Representations, December 2024
Yucheng Hu, Yanjiang Guo, Pengchao Wang, Xiaoyu Chen, Yen-Jen Wang, Jianke Zhang, Koushil Sreenath, Chaochao Lu, and Jianyu Chen. Video Prediction Pol- icy: A Generalist Robot Policy with Predictive Visual Representations, December 2024
2024
-
[39]
Goal-Guided Transformer-Enabled Reinforcement Learning for Efficient Autonomous Navigation
Wenhui Huang, Yanxin Zhou, Xiangkun He, and Chen Lv. Goal-Guided Transformer-Enabled Reinforcement Learning for Efficient Autonomous Navigation. IEEE Transactions on Intelligent Transportation Systems , 25 (2):1832–1845, February 2024. ISSN 1558-0016. doi: 10.1109/TITS.2023.3312453
2024
-
[40]
Rellis-3d dataset: Data, benchmarks and anal- ysis
Peng Jiang, Philip Osteen, Maggie Wigness, and Srikanth Saripalli. Rellis-3d dataset: Data, benchmarks and anal- ysis. In 2021 IEEE international conference on robotics and automation (ICRA) , pages 1110–1116. IEEE, 2021
2021
-
[41]
Johnson, Uday S
Jacob J. Johnson, Uday S. Kalra, Ankit Bhatia, Linjun Li, Ahmed H. Qureshi, and Michael C. Yip. Motion Planning Transformers: A Motion Planning Framework for Mobile Robots, November 2022
2022
-
[42]
Beyond Sight: Finetuning Generalist Robot Policies with Hetero- geneous Sensors via Language Grounding, January 2025
Joshua Jones, Oier Mees, Carmelo Sferrazza, Kyle Sta- chowicz, Pieter Abbeel, and Sergey Levine. Beyond Sight: Finetuning Generalist Robot Policies with Hetero- geneous Sensors via Language Grounding, January 2025
2025
-
[43]
VI-IKD: High-speed ac- curate off-road navigation using learned visual-inertial inverse kinodynamics
Haresh Karnan, Kavan Singh Sikand, Pranav Atreya, Sadegh Rabiee, Xuesu Xiao, Garrett Warnell, Peter Stone, and Joydeep Biswas. VI-IKD: High-speed ac- curate off-road navigation using learned visual-inertial inverse kinodynamics. In 2022 IEEE/RSJ International Conference on Int...
2022
-
[44]
DINO-Foresight Looking into the Future with DINO, December 2024
Efstathios Karypidis, Ioannis Kakogeorgiou, Spyros Gi- daris, and Nikos Komodakis. DINO-Foresight Looking into the Future with DINO, December 2024
2024
-
[45]
MarineFormer: A Spatio-Temporal Attention Model for USV Navigation in Dynamic Marine Environments, December 2024
Ehsan Kazemi and Iman Soltani. MarineFormer: A Spatio-Temporal Attention Model for USV Navigation in Dynamic Marine Environments, December 2024
2024
-
[46]
Daniel Lawson and Ahmed H. Qureshi. Control Transformer: Robot Navigation in Unknown Environ- ments Through PRM-Guided Return-Conditioned Se- quence Modeling. In 2023 IEEE/RSJ International Con- ference on Intelligent Robots and Systems (IROS) , pages 9324–9331, October 2023. ...
2023
-
[47]
Learning terrain-aware kinodynamic model for autonomous off-road rally driving with model predictive path integral control
Hojin Lee, Taekyung Kim, Jungwi Mun, and Wonsuk Lee. Learning terrain-aware kinodynamic model for autonomous off-road rally driving with model predictive path integral control. IEEE Robotics and Automation Letters, 2023
2023
-
[48]
CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos, November 2024
Xinhao Liu, Jintong Li, Yicheng Jiang, Niranjan Su- jay, Zhicheng Yang, Juexiao Zhang, John Abanes, Jing Zhang, and Chen Feng. CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos, November 2024
2024
-
[49]
Efficient training of visual transformers with small datasets
Yahui Liu, Enver Sangineto, Wei Bi, Nicu Sebe, Bruno Lepri, and Marco Nadai. Efficient training of visual transformers with small datasets. Advances in Neural Information Processing Systems, 34:23818–23830, 2021
2021
-
[50]
Decoupled Weight Decay Regularization, January 2019
Ilya Loshchilov and Frank Hutter. Decoupled Weight Decay Regularization, January 2019
2019
-
[51]
nGPT: Normalized Transformer with Representation Learning on the Hypersphere, October 2024
Ilya Loshchilov, Cheng-Ping Hsieh, Simeng Sun, and Boris Ginsburg. nGPT: Normalized Transformer with Representation Learning on the Hypersphere, October 2024
2024
-
[52]
Discrete Rep- resentations Strengthen Vision Transformer Robustness, April 2022
Chengzhi Mao, Lu Jiang, Mostafa Dehghani, Carl V on- drick, Rahul Sukthankar, and Irfan Essa. Discrete Rep- resentations Strengthen Vision Transformer Robustness, April 2022
2022
-
[53]
GPT-Driver: Learning to Drive with GPT, October 2023
Jiageng Mao, Yuxi Qian, Hang Zhao, and Yue Wang. GPT-Driver: Learning to Drive with GPT, October 2023
2023
-
[54]
Wild Visual Navigation: Fast Traversability Learning via Pre-Trained Models and On- line Self-Supervision, April 2024
Mat ´ıas Mattamala, Jonas Frey, Piotr Libera, Nived Che- brolu, Georg Martius, Cesar Cadena, Marco Hutter, and Maurice Fallon. Wild Visual Navigation: Fast Traversability Learning via Pre-Trained Models and On- line Self-Supervision, April 2024
2024
-
[55]
Elevation mapping for locomotion and navigation using gpu
Takahiro Miki, Lorenz Wellhausen, Ruben Grandia, Fabian Jenelten, Timon Homberger, and Marco Hutter. Elevation mapping for locomotion and navigation using gpu. In 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 2273–2280. IEEE, 2022
2022
-
[56]
VertiEncoder: Self-Supervised Kinodynamic Representation Learning on Vertically Challenging Terrain, September 2024
Mohammad Nazeri, Aniket Datar, Anuj Pokhrel, Chenhui Pan, Garrett Warnell, and Xuesu Xiao. VertiEncoder: Self-Supervised Kinodynamic Representation Learning on Vertically Challenging Terrain, September 2024
2024
-
[57]
V ANP: Learning Where to See for Navigation with Self-Supervised Vision-Action Pre- Training
Mohammad Nazeri, Junzhe Wang, Amirreza Payandeh, and Xuesu Xiao. V ANP: Learning Where to See for Navigation with Self-Supervised Vision-Action Pre- Training. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 2741–2746, Abu Dhabi, United...
2024
-
[58]
Ex- ploring Reflective Limitation of Behavior Cloning in Autonomous Vehicles
Mohammad Hossein Nazeri and Mahdi Bohlouli. Ex- ploring Reflective Limitation of Behavior Cloning in Autonomous Vehicles. In 2021 IEEE International Conference on Data Mining (ICDM) , pages 1252–1257, December 2021. doi: 10.1109/ICDM51629.2021.00153
2021
-
[59]
Octo: An Open-Source Generalist Robot Policy, May 2024
Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Tobias Kreiman, Charles Xu, Jianlan Luo, You Liang Tan, Pannag Sanketi, Quan Vuong, Ted Xiao, Dorsa Sadigh, Chelsea Finn, and Sergey Levine. Octo: An Open-Source General...
2024
-
[60]
DINOv2: Learning Robust Visual Features without Supervision, April 2023
Maxime Oquab, Timoth ´ee Darcet, Th ´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fer- nandez, Daniel Haziza, Francisco Massa, Alaaeldin El- Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Micha...
2023
-
[61]
Fast local planning and mapping in unknown off-road terrain
Timothy Overbye and Srikanth Saripalli. Fast local planning and mapping in unknown off-road terrain. In 2020 IEEE International Conference on Robotics and Automation (ICRA), pages 5912–5918. IEEE, 2020
2020
-
[62]
Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration
Abby O’Neill, Abdul Rehman, Abhiram Maddukuri, Ab- hishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, et al. Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration. In 2024 IEEE Internat...
2024
-
[63]
Tra- verse the Non-Traversable: Estimating Traversability for Wheeled Mobility on Vertically Challenging Terrain, September 2024
Chenhui Pan, Aniket Datar, Anuj Pokhrel, Matthew Choulas, Mohammad Nazeri, and Xuesu Xiao. Tra- verse the Non-Traversable: Estimating Traversability for Wheeled Mobility on Vertically Challenging Terrain, September 2024
2024
-
[64]
Imitation learning for agile autonomous driving
Yunpeng Pan, Ching-An Cheng, Kamil Saigol, Keuntaek Lee, Xinyan Yan, Evangelos A Theodorou, and Byron Boots. Imitation learning for agile autonomous driving. The International Journal of Robotics Research , 2020
2020
-
[65]
Viorica P ˘atr˘aucean, Xu Owen He, Joseph Heyward, Chuhan Zhang, Mehdi S. M. Sajjadi, George-Cristian Muraru, Artem Zholus, Mahdi Karami, Ross Goroshin, Yutian Chen, Simon Osindero, Jo˜ao Carreira, and Razvan Pascanu. TRecViT: A Recurrent Video Transformer, December 2024
2024
-
[66]
Transformers for Image-Goal Naviga- tion, May 2024
Nikhilanj Pelluri. Transformers for Image-Goal Naviga- tion, May 2024
2024
-
[67]
FAST: Efficient Action To- kenization for Vision-Language-Action Models, January 2025
Karl Pertsch, Kyle Stachowicz, Brian Ichter, Danny Driess, Suraj Nair, Quan Vuong, Oier Mees, Chelsea Finn, and Sergey Levine. FAST: Efficient Action To- kenization for Vision-Language-Action Models, January 2025
2025
-
[68]
CAHSOR: Competence-aware high-speed off-road ground navigation in SE (3)
Anuj Pokhrel, Aniket Datar, Mohammad Nazeri, and Xuesu Xiao. CAHSOR: Competence-aware high-speed off-road ground navigation in SE (3). IEEE Robotics and Automation Letters, 9(11):9653–9660, 2024
2024
-
[69]
Pomerleau
Dean A. Pomerleau. ALVINN: An autonomous land vehicle in a neural network. In D. Touretzky, editor, Advances in Neural Information Processing Systems , volume 1. Morgan-Kaufmann, 1988
1988
-
[70]
Improving language understanding by generative pre-training
Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training. 2018
2018
-
[71]
Language models are unsupervised multitask learners
Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. 2019
2019
-
[72]
An Empirical Study of Autoregressive Pre-training from Videos, January 2025
Jathushan Rajasegaran, Ilija Radosavovic, Rahul Ravis- hankar, Yossi Gandelsman, Christoph Feichtenhofer, and Jitendra Malik. An Empirical Study of Autoregressive Pre-training from Videos, January 2025
2025
-
[73]
Neural Machine Translation of Rare Words with Sub- word Units, June 2016
Rico Sennrich, Barry Haddow, and Alexandra Birch. Neural Machine Translation of Rare Words with Sub- word Units, June 2016
2016
-
[74]
Learning Off-Road Terrain Traversability with Self-Supervisions Only, May 2023
Junwon Seo, Sungdae Sim, and Inwook Shim. Learning Off-Road Terrain Traversability with Self-Supervisions Only, May 2023
2023
-
[75]
Masked World Models for Visual Control, May 2023
Younggyo Seo, Danijar Hafner, Hao Liu, Fangchen Liu, Stephen James, Kimin Lee, and Pieter Abbeel. Masked World Models for Visual Control, May 2023
2023
-
[76]
Multi-View Masked World Models for Visual Robotic Manipulation, February 2023
Younggyo Seo, Junsu Kim, Stephen James, Kimin Lee, Jinwoo Shin, and Pieter Abbeel. Multi-View Masked World Models for Visual Robotic Manipulation, February 2023
2023
-
[77]
Ramp: A risk- aware mapping and planning pipeline for fast off-road ground robot navigation
Lakshay Sharma, Michael Everett, Donggun Lee, Xiaoyi Cai, Philip Osteen, and Jonathan P How. Ramp: A risk- aware mapping and planning pipeline for fast off-road ground robot navigation. In 2023 IEEE International Conference on Robotics and Automation (ICRA) , pages 5730–5736. ...
2023
-
[78]
Robot adaptation to unstructured terrains by joint representation and apprenticeship learning
Sriram Siva, Maggie Wigness, John Rogers, and Hao Zhang. Robot adaptation to unstructured terrains by joint representation and apprenticeship learning. In Robotics: Science and Systems (RSS) , 2019
2019
-
[79]
Nauts: Negotiation for adapta- tion to unstructured terrain surfaces
Sriram Siva, Maggie Wigness, John G Rogers, Long Quang, and Hao Zhang. Nauts: Negotiation for adapta- tion to unstructured terrain surfaces. In 2022 IEEE/RSJ International Conference on Intelligent Robots and Sys- tems (IROS), pages 1733–1740. IEEE, 2022
2022
-
[80]
Improving off-road planning techniques with learned costs from physical interactions
Matthew Sivaprakasam, Samuel Triest, Wenshan Wang, Peng Yin, and Sebastian Scherer. Improving off-road planning techniques with learned costs from physical interactions. In 2021 IEEE International Conference on Robotics and Automation (ICRA) , pages 4844–4850. IEEE, 2021
2021
-
[81]
How to train your ViT? Data, augmentation, and regular- ization in vision transformers
Andreas Peter Steiner, Alexander Kolesnikov, Xiaohua Zhai, Ross Wightman, Jakob Uszkoreit, and Lucas Beyer. How to train your ViT? Data, augmentation, and regular- ization in vision transformers. Transactions on Machine Learning Research, 2022. ISSN 2835-8856
2022
-
[82]
Scalabil- ity in perception for autonomous driving: Waymo open dataset
Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aure- lien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al. Scalabil- ity in perception for autonomous driving: Waymo open dataset. In Proceedings of the IEEE/CVF conference on compute...
2020
-
[83]
Tartandrive: A large-scale dataset for learning off-road dynamics models
Samuel Triest, Matthew Sivaprakasam, Sean J Wang, Wenshan Wang, Aaron M Johnson, and Sebastian Scherer. Tartandrive: A large-scale dataset for learning off-road dynamics models. In 2022 International Confer- ence on Robotics and Automation (ICRA) , pages 2546–
2022
-
[84]
Attention is all you need
Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems , 30, 2017
2017
-
[85]
Nav- Former: A Transformer Architecture for Robot Target- Driven Navigation in Unknown and Dynamic Environ- ments
Haitong Wang, Aaron Hao Tan, and Goldie Nejat. Nav- Former: A Transformer Architecture for Robot Target- Driven Navigation in Unknown and Dynamic Environ- ments. 2024. doi: 10.48550/ARXIV .2402.06838
-
[86]
Navigating the Landscape of Large Lan- guage Models: A Comprehensive Review and Analysis of Paradigms and Fine-Tuning Strategies, April 2024
Benjue Weng. Navigating the Landscape of Large Lan- guage Models: A Comprehensive Review and Analysis of Paradigms and Fine-Tuning Strategies, April 2024
2024
-
[87]
A rugd dataset for autonomous navigation and visual perception in unstructured outdoor environments
Maggie Wigness, Sungmin Eum, John G Rogers, David Han, and Heesung Kwon. A rugd dataset for autonomous navigation and visual perception in unstructured outdoor environments. In 2019 IEEE/RSJ International Confer- ence on Intelligent Robots and Systems (IROS) , pages 5000–5007....
2019
-
[88]
Model predictive path integral control: From theory to parallel computation
Grady Williams, Andrew Aldrich, and Evangelos A Theodorou. Model predictive path integral control: From theory to parallel computation. Journal of Guidance, Control, and Dynamics , 2017
2017
-
[89]
Dolan, and Guanya Shi
Wenli Xiao, Haoru Xue, Tony Tao, Dvij Kalaria, John M. Dolan, and Guanya Shi. AnyCar to Anywhere: Learn- ing Universal Dynamics Model for Agile and Adaptive Mobility, September 2024
2024
-
[90]
Learning inverse kinodynamics for accurate high-speed off-road navigation on unstructured terrain
Xuesu Xiao, Joydeep Biswas, and Peter Stone. Learning inverse kinodynamics for accurate high-speed off-road navigation on unstructured terrain. IEEE Robotics and Automation Letters, 6(3):6054–6060, 2021
2021
-
[91]
Motion planning and control for mobile robot navigation using machine learning: A survey
Xuesu Xiao, Bo Liu, Garrett Warnell, and Peter Stone. Motion planning and control for mobile robot navigation using machine learning: A survey. Autonomous Robots, 46(5):569–597, June 2022. ISSN 0929-5593, 1573-7527. doi: 10.1007/s10514-022-10039-8
2022 doi
-
[92]
On layer normalization in the transformer architecture
Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu. On layer normalization in the transformer architecture. In International Conference on Machine Learning, pages 10524–10533. PMLR, 2020. doi: 10.5555/...
2020
-
[93]
Prince, and Yanshuai Cao
Peng Xu, Dhruv Kumar, Wei Yang, Wenjie Zi, Keyi Tang, Chenyang Huang, Jackie Chi Kit Cheung, Si- mon J.D. Prince, and Yanshuai Cao. Optimizing Deeper Transformers on Small Datasets. In Proceedings of the 59th Annual Meeting of the Association for Com- putational Linguistics an...
2021 doi
-
[94]
Scaling vision transformers
Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer. Scaling vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 12104–12113, June 2022
2022
-
[95]
Root mean square layer normalization
Biao Zhang and Rico Sennrich. Root mean square layer normalization. In Advances in Neural Information Processing Systems 32 , Vancouver, Canada, 2019
2019
-
[96]
NaviFormer: A Data- Driven Robot Navigation Approach via Sequence Mod- eling and Path Planning with Safety Verification
Xuyang Zhang, Ziyang Feng, Quecheng Qiu, Yu’an Chen, Bei Hua, and Jianmin Ji. NaviFormer: A Data- Driven Robot Navigation Approach via Sequence Mod- eling and Path Planning with Safety Verification. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages...
2024
-
[97]
DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning, November 2024
Gaoyue Zhou, Hengkai Pan, Yann LeCun, and Lerrel Pinto. DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning, November 2024. APPENDIX A MODEL ARCHITECTURE TABLE II: V ERTI FORMER Architecture Parameters. VERTI ENCODER Layers 6 Normalization RMSNorm [9...
2024
-
[1703]
PMLR, 2020-07-13/2020-07-18
2020
Reviewed August 9, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.