Pith. sign in

REVIEW 4 major objections 4 minor 1 cited by

VertiFormer: A Data-Efficient Multi-Task Transformer for Off-Road Robot Mobility

T0 review · 4 major / 4 minor · reviewed 2026-08-09 · deepseek-v4-flash

Pith's one-line read VertiFormer predicts next pose, action, and terrain patch from one hour of driving data, and beats a terrain-specialist kinodynamic model on a physical robot.

desk verdict A genuinely new multi-task transformer for off-road kinodynamics, with a load-bearing data-claim contradiction that must be fixed before the efficiency story holds. read the letter →

arxiv 2502.00543 v1 pith:QDJ7KOVT submitted 2025-02-01 cs.RO cs.CVcs.LG

classification cs.ROcs.CVcs.LG
keywords off-roadmobilityTransformerkinodynamicmodelingmulti-tasklearningmaskedbehaviorcloningnon-autoregressivepredictiondata-efficient
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper sets out to show that a Transformer can learn enough about vehicle-terrain interaction to drive a wheeled robot over rough, rocky terrain using only one hour of teleoperated demonstrations. It proposes VertiFormer, which maps past poses, actions, and terrain patches into a single latent sequence and trains with a learnable mask so that one model can predict the next pose, the next action, or both. The authors argue that this unified representation, combined with non-autoregressive prediction, makes Transformers data-efficient for mobility, in contrast to NLP and CV practice that typically relies on internet-scale data. If correct, off-road robots could obtain forward kinodynamics, inverse kinodynamics, and behavior cloning from a single short data collection and run the model onboard.

What carries the argument

The central object is the unified multi-modal latent token $z_t = f_s(\hat a_t \cdot \hat p_t \cdot \hat i_t)$, a learned linear projection that concatenates embedded actions, poses, and terrain patches into one homogeneous token stream before the encoder. This shared token space is the inductive bias that the paper credits for letting one hour of data carry three tasks. Training alternates two learned masks with equal probability: action-conditioned pose prediction for forward kinodynamics, and pose-conditioned action prediction for inverse kinodynamics, while masking both enables zero-shot behavior cloning. Multi-context tokens and cross-attention with causal masking let the decoder emit all future steps non-autoregressively, avoiding the error accumulation that autoregressive decoding suffers over long horizons.

What would settle it

Train the same multi-task masking and non-autoregressive decoder with separate per-modality tokens instead of the unified latent projection, and compare held-out forward, inverse, and behavior-cloning error on the same one-hour dataset. If error rates are statistically indistinguishable, the unified representation is not the cause of the data-efficiency gain. Also, count how many rollover and high-roll episodes occur in the one-hour log; if those states appear only a handful of times, the claim that the log covers dangerous kinodynamic interactions is weak.

Watch

Extended reading notes

Core claim

On the paper's own terms, the discovery is that a Transformer can be made data-efficient for off-road mobility by replacing separate per-modality tokens with a unified latent representation and by alternating learnable masks between future actions and future poses during training. Trained on one hour of human teleoperation over a rock testbed, VertiFormer predicts the next pose, action, and terrain patch simultaneously, and its forward kinodynamic model, when plugged into a sampling-based planner, completes ten of ten traversals on an unseen terrain testbed while a state-of-the-art kinodynamic model specialized for vertically challenging terrain completes eight of ten. The same architecture also produces inverse kinodynamics and zero-shot behavior cloning without retraining dedicated heads.

Load-bearing premise

The load-bearing premise is that combining pose, action, and terrain into one latent token makes all three tasks share a representation that one hour of data can learn; the paper states this as a hypothesis and supports it mainly with a temporal-order classification ablation rather than direct downstream-task evidence.

Editorial extensions

If this is right

  • A single one-hour teleoperation log can be enough to train a forward kinodynamic model that sampling-based planners can use to traverse rocky, vertically challenging terrain.
  • The same trained weights handle inverse kinodynamics and zero-shot behavior cloning, so a robot can switch tasks or tolerate missing future pose or action inputs without retraining separate heads.
  • Because the decoder is non-autoregressive, one-second and two-second predictions drift less than an autoregressive decoder's, which matters for real-time planning loops.
  • Under extreme data scarcity, sinusoidal positional encoding and a final normalization layer help, while an auxiliary terrain-patch reconstruction head can degrade the primary tasks.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • I infer that the equal 50/50 split between action-conditioned and pose-conditioned masking is a design choice, not a tuned optimum; varying this ratio and the mask token's learning rate could change how quickly shared representations emerge from small data.
  • I infer that the unified-token projection could transfer to other locomotion forms, such as legged robots or tracked vehicles, whenever states and commands admit vector embeddings; the paper only demonstrates wheeled motion.
  • I infer that the temporal-order classification result is a proxy, not a measure of downstream task quality; a cleaner test would ablate unified versus separate tokens and report forward, inverse, and cloning error on held-out terrain.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, and a circularity audit.

Referee Report

4 major / 4 minor

Summary. VertiFormer is a Transformer architecture for off-road mobility that combines a Transformer encoder and decoder, a unified latent representation for actions, poses, and terrain patches, learnable masking, and non-autoregressive multi-token prediction. It is trained on human teleoperation data claimed to be one hour long and evaluated on forward and inverse kinodynamic modeling and behavior cloning, first as offline prediction error against TAL and VertiEncoder, then on a physical V4W robot with MPPI, Dijkstra, and behavior-cloning controllers. The paper reports ablations on positional encoding, normalization, unified representation, patch head, prediction horizon, and training paradigm. The stated contributions are a data-efficient multi-task Transformer, empirical design guidelines, and physical-robot demonstrations.

Significance. If the central claims held, this would be a practically relevant demonstration that a Transformer can learn multiple off-road mobility tasks from roughly one hour of demonstration data and run onboard a robot, with a released open-source implementation. The strengths of the paper are the real-robot evaluation, the systematic ablation structure, and the clear architectural description. However, the headline data-efficiency claim is currently undercut by the dataset description in Appendix B, and the numerical comparisons lack error bars, confidence intervals, or significance tests. The evidence for the central multi-task mechanism is indirect. These issues are fixable but currently prevent the paper's claims from being fully verified.

major comments (4)
  1. [Appendix B vs. Abstract/§IV] The paper repeatedly states that VertiFormer is 'trained with only one hour of data' (Abstract, §I, §IV, §IV-C), but Appendix B, the only dataset description, says: 'The dataset includes 30 minutes of data from both a planar surface and the rock testbed.' This is ambiguous between 30 minutes total and 30 minutes per surface, but in either reading it is not a statement of one hour. Because the data-efficiency claim and the comparison with TAL [19] depend on the exact training-set size and composition, the authors must reconcile this discrepancy, report the exact duration, the proportion of planar vs. rock-teleoperation data, and the precise train/test split, and state whether TAL and all other baselines were trained on the identical data. Without this audit, the headline result cannot be evaluated.
  2. [§III-B and Table I] The headline offline comparison reports error rates of 0.495 (VertiFormer), 0.516 (Nazeri et al.), and 0.528 (TAL) with no confidence intervals, number of test trajectories, or significance tests; the 0.033 difference between VertiFormer and TAL is small and could be sampling noise. Similarly, Table I reports success rates out of 10 physical trials (e.g., FKD: TAL 8/10 vs. VertiFormer 10/10) without confidence intervals, per-trial distributions, or a statistical test, and the mean roll/pitch columns include wide standard deviations. The claim that VertiFormer 'outperforms' TAL on navigation performance is therefore not statistically established. Please report repeated-seed or bootstrap intervals for the offline metrics, per-trial results for Table I, and an appropriate significance test or an explicit statement that the differences are descriptive only.
  3. [§IV-B / Fig. 5 and §III-A2] The only direct evidence offered for the unified-latent-representation benefit is Fig. 5, an auxiliary temporal-order classification loss, shown as a single learning curve with no error bars and no downstream FKD/IKD/BC results. The paper itself labels the multi-task data-efficiency mechanism as 'hypothesized' in §III-A2. To support the central claim that alternating masking plus unified representation drives the data-efficiency gain, the authors should show the unified vs. separate-modality comparison on the actual downstream tasks and include variance over training seeds. As it stands, the paper does not establish that the gains come from multi-task sharing rather than from the unified input representation alone.
  4. [§IV-C / Fig. 7 and Abstract] The abstract and introduction state that VertiFormer predicts the next pose, action, and terrain patch, but Fig. 7 shows that adding the patch reconstruction head degrades performance on the primary tasks, and the text explains that the patch head 'introduces noise into the learning process.' It is never stated whether the final deployed VertiFormer includes the patch head. If it does not, the terrain-patch prediction contribution is not part of the final model; if it does, the ablation appears to contradict the design. The authors should state the final architecture explicitly, including which heads are used in the robot experiments, and align the abstract and introduction with that choice.
minor comments (4)
  1. [Throughout] The notation 'V ERTI FORMER', 'V ERTI ENCODER', and 'V ERTI DECODER' is typeset with irregular spacing; use a consistent single-token name throughout.
  2. [Appendix A, Table III] Table III cites [73] (Sennrich et al., subword units) for batch normalization; this appears to be the wrong reference and should be corrected.
  3. [§III-B] The offline error-rate table has no table number or caption; add one and specify how the error rate is computed (e.g., normalized by what quantity) and over how many held-out samples.
  4. [Fig. 6] Fig. 6 includes a panel labeled 'IKD' alongside X/Y/Z panels, but the IKD metric is not defined in the caption; clarify what is plotted in each panel and the units.

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity: the reported predictions are evaluated on held-out data, and the self-citations are empirical comparators rather than load-bearing derivation inputs.

full rationale

The paper's derivation chain is self-contained with respect to circularity. VertiFormer's FKD, IKD, and BC outputs are obtained by training the model on a fixed dataset and measuring prediction error against ground truth, with no parameter fitted directly to the reported success metrics and then renamed as a prediction. The unified latent representation, learnable masking, and non-autoregressive training are described as architectural design choices, and the multi-task benefit is explicitly labeled a hypothesis in Section III-A2 rather than being assumed through an equation that defines the output in terms of the input. The baselines TAL [19] and VertiEncoder [56] are prior works with overlapping authorship, but they are used as experimental comparators for physical-robot outcomes, not as justifications for why VertiFormer's architecture must work; no uniqueness theorem or self-citation chain forces the model choice. The paper does contain a notable internal inconsistency: the main text repeatedly claims 'one hour' of training data, while Appendix B states 'The dataset includes 30 minutes of data from both a planar surface and the rock testbed.' This is a reporting and audit concern about the size of the training set, and it may undermine the quantitative force of the data-efficiency claim, but it is not a circular reduction of a prediction to its inputs. No equation in the paper equates a claimed result with a fitted parameter or defines an output in terms of the target quantity.

Assumptions & free parameters 7 free parameters · 4 assumptions · 0 invented entities

The paper's central claims rest on architecture design choices, a specific data collection protocol, and an unverified assumption that multi-task masking drives the data-efficiency gains. No new physical entities or forces are introduced, so the invented-entities ledger is empty.

free parameters (7)
  • hidden size D = 512
    Chosen by hand; no sensitivity analysis is reported.
  • encoder/decoder layers = 6 encoder, 4 decoder
    Chosen by hand; architecture table in Appendix A.
  • attention heads = 8
    Chosen by hand; Appendix A.
  • dropout = 0.3
    Chosen by hand; Appendix A.
  • learning rate and weight decay = 5e-4, 0.08
    AdamW settings in Appendix B, chosen without reported tuning curves.
  • batch size and epochs = 512, 200
    Training schedule in Appendix B.
  • prediction horizon tau = 3 steps at 3 Hz
    Future context is fixed at three steps; the paper notes in Limitations that changing horizon requires retraining.
assumptions (4)
  • domain assumption Transformer self-attention can learn kinodynamic mappings from sequential robot data when inputs are projected into a unified embedding.
    This is the central architectural bet; evidence is empirical only, mainly Fig. 5.
  • domain assumption Elevation-map terrain patches and 6-DoF odometry contain sufficient information to model tire deformation, suspension travel, and wheel-terrain interaction.
    No physics-based validation or sensor noise analysis is provided.
  • domain assumption Human teleoperation on the rock testbed is a valid expert policy for training FKD, IKD, and BC.
    All three objectives are defined from this demonstration data.
  • ad hoc to paper Alternating action and pose masking improves data efficiency by forcing shared multi-task latent representations.
    The paper calls this 'hypothesized' in Section III-A2 and supports it with a single order-classification ablation.

how reviews work

0 comments
Cite this review

Pith. "Pith review of VertiFormer: A Data-Efficient Multi-Task Transformer for Off-Road Robot Mobility." pith.science (2026). https://pith.science/paper/QDJ7KOVT

@misc{pith2026250200543,
  author       = {Pith},
  title        = {Pith review of: VertiFormer: A Data-Efficient Multi-Task Transformer for Off-Road Robot Mobility},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/QDJ7KOVT}},
  note         = {Machine review of arXiv:2502.00543}
}
read the original abstract

Sophisticated learning architectures, e.g., Transformers, present a unique opportunity for robots to understand complex vehicle-terrain kinodynamic interactions for off-road mobility. While internet-scale data are available for Natural Language Processing (NLP) and Computer Vision (CV) tasks to train Transformers, real-world mobility data are difficult to acquire with physical robots navigating off-road terrain. Furthermore, training techniques specifically designed to process text and image data in NLP and CV may not apply to robot mobility. In this paper, we propose VertiFormer, a novel data-efficient multi-task Transformer model trained with only one hour of data to address such challenges of applying Transformer architectures for robot mobility on extremely rugged, vertically challenging, off-road terrain. Specifically, VertiFormer employs a new learnable masked modeling and next token prediction paradigm to predict the next pose, action, and terrain patch to enable a variety of off-road mobility tasks simultaneously, e.g., forward and inverse kinodynamics modeling. The non-autoregressive design mitigates computational bottlenecks and error propagation associated with autoregressive models. VertiFormer's unified modality representation also enhances learning of diverse temporal mappings and state representations, which, combined with multiple objective functions, further improves model generalization. Our experiments offer insights into effectively utilizing Transformers for off-road robot mobility with limited data and demonstrate our efficiently trained Transformer can facilitate multiple off-road mobility tasks onboard a physical mobile robot.

Figures

Figures reproduced from arXiv: 2502.00543 by the authors.

Figure 1
Figure 1. VERTIFORMER is a data-efficient multi-task Transformer specifically for off-road mobility. Leveraging kinodynamic representation learning, VERTIFORMER employs unified multi-modal latent representation, learnable masked modeling, and non-autoregressive training to understand complex and nuanced vehicle-terrain interactions with only one hour of training data. Abstract—Sophisticated learning architectures, e.g., Trans… view at source ↗
Figure 2
Figure 2. VERTIFORMER Architecture. VERTIFORMER employs a TransformerEncoder (left) to receive a history of terrain patches, actions, and poses along with multiple context tokens. To predict future states, the model computes cross-attention between these context tokens and the masked upcoming actions or poses. Causal masking is implemented during this cross￾attention computation to ensure that predictions are conditioned only… view at source ↗
Figure 3
Figure 3. Positional Encoding: Sinusoidal positional encoding achieves better model accuracy than learnable encoding for predicting X, Y, and Z components of the robot pose. X Y Z 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 Error Rate in 1s 0.218 0.385 0.763 0.210 0.360 0.754 Without Last Normalization With Last Normalization [PITH_FULL_IMAGE:figures/full_fig_p005_3.png] view at source ↗
Figures from the paper (9 more)
Figure 4
Figure 4. Figure 4: Normalizing Output: Normalizing the Trans￾former output before passing the embeddings to the task decoder improves model performance. enable the model to effectively process sequential data. Learn￾able positional encodings, typically implemented as trainable vectors ad…
Figure 5
Figure 5. Figure 5: Kinodynamics Understanding: Without unified latent representation the model cannot capture temporal dependen￾cies and understand kinodynamic transitions, resulting in an almost flat learning curve. Unified latent space representation offers a significant advantage in s…
Figure 6
Figure 6. Figure 6: Prediction Horizon: VERTIFORMER is capable of predicting a longer horizon without losing much accuracy due to its non-autoregressive nature. Prediction horizon is a critical factor in navigation plan￾ning. While longer prediction horizons can potentially lead to better…
Figure 8
Figure 8. Figure 8: MM vs NTP vs End2End: VERTIFORMER achieves best accuracy across FKD, IKD, and BC compared to VER￾TIENCODER (MM), VERTIDECODER (NTP), and End2End. autoregressive NTP (VERTIDECODER, [PITH_FULL_IMAGE:figures/full_fig_p007_8.png]
Figure 7
Figure 7. Figure 7: Patch Prediction Head: The inclusion of a patch reconstruction head results in a degradation of overall model performance. This counterintuitive result can be attributed to the inherent difficulty in accurately predicting the detailed structure of off-road terrain topo…
Figure 9
Figure 9. Figure 9: Unseen Test Environments with Rocks/Boulders, [PITH_FULL_IMAGE:figures/full_fig_p008_9.png]
Figure 10
Figure 10. Figure 10: Qualitative Results of 3-Step and 6-Step Successful and Failed Trajectory Prediction over One and Two Second(s). [PITH_FULL_IMAGE:figures/full_fig_p016_10.png]
Figure 11
Figure 11. Figure 11: Qualitative Comparison of Drifting between Non-Autoregressive V [PITH_FULL_IMAGE:figures/full_fig_p017_11.png]
Figure 12
Figure 12. Figure 12: Visualization of VERTIFORMER Predictions in green and Ground Truth in blue [PITH_FULL_IMAGE:figures/full_fig_p017_12.png]

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 1 Pith paper

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Narrate2Nav: Real-Time Visual Navigation with Implicit Language Reasoning in Human-Centric Environments

    cs.RO 2025-06 conditional novelty 6.0 of 10

    Narrate2Nav uses Barlow Twins alignment to distill language-based reasoning from a large teacher into a small RGB-only navigation model, reporting lower trajectory error and higher goal-reaching success than four baselines.

Reference graph

Works this paper leans on

97 extracted references · 68 canonical work pages · cited by 1 Pith paper

  1. [19]

    Terrain-Attentive Learning for Efficient 6-DoF Kinodynamic Modeling on Vertically Challenging Terrain

    Aniket Datar, Chenhui Pan, Mohammad Nazeri, Anuj Pokhrel, and Xuesu Xiao. Terrain-Attentive Learning for Efficient 6-DoF Kinodynamic Modeling on Vertically Challenging Terrain. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS) , pages 5438–5443, Abu Dhabi, United Arab Emirates, October 2024. IEEE. ISBN 979-8-3503-7770-5. d...

  2. [1]

    Invariance is Key to Generalization: Examining the Role of Representation in Sim-to-Real Transfer for Visual Navigation, December 2023

    Bo Ai, Zhanxin Wu, and David Hsu. Invariance is Key to Generalization: Examining the Role of Representation in Sim-to-Real Transfer for Visual Navigation, December 2023

  3. [2]

    Self-supervised learning from images with a joint-embedding predictive architecture

    Mahmoud Assran, Quentin Duval, Ishan Misra, Piotr Bo- janowski, Pascal Vincent, Michael Rabbat, Yann LeCun, and Nicolas Ballas. Self-supervised learning from images with a joint-embedding predictive architecture. arXiv preprint arXiv:2301.08243, 2023

  4. [3]

    Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E. Hinton. Layer Normalization, July 2016

  5. [4]

    Curriculum learning for vehicle lateral stability estimations

    Jihwan Bae, Taekyung Kim, Wonsuk Lee, and Inwook Shim. Curriculum learning for vehicle lateral stability estimations. IEEE Access, 9:89249–89262, 2021

  6. [5]

    Navigation World Models, December 2024

    Amir Bar, Gaoyue Zhou, Danny Tran, Trevor Darrell, and Yann LeCun. Navigation World Models, December 2024

  7. [6]

    MC- JEPA: A Joint-Embedding Predictive Architecture for Self-Supervised Learning of Motion and Content Fea- tures, July 2023

    Adrien Bardes, Jean Ponce, and Yann LeCun. MC- JEPA: A Joint-Embedding Predictive Architecture for Self-Supervised Learning of Motion and Content Fea- tures, July 2023

  8. [7]

    Revisiting Feature Prediction for Learning Visual Representations from Video, February 2024

    Adrien Bardes, Quentin Garrido, Jean Ponce, Xinlei Chen, Michael Rabbat, Yann LeCun, Mahmoud Assran, and Nicolas Ballas. Revisiting Feature Prediction for Learning Visual Representations from Video, February 2024

Show all 97 references
  1. [8]

    Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, Xin Zhang, Jake Zhao, and Karol Zieba

    Mariusz Bojarski, Davide Del Testa, Daniel Dworakowski, Bernhard Firner, Beat Flepp, Prasoon Goyal, Lawrence D. Jackel, Mathew Monfort, Urs Muller, Jiakai Zhang, Xin Zhang, Jake Zhao, and Karol Zieba. End to End Learning for Self-Driving Cars, April 2016

  2. [9]

    A survey on terrain traversability analysis for autonomous ground ve- hicles: Methods, sensors, and challenges

    Paulo Borges, Thierry Peynot, Sisi Liang, Bilal Arain, Matt Wildie, Melih Minareci, Serge Lichman, Garima Samvedi, Inkyu Sa, Nicolas Hudson, Michael Milford, Peyman Moghadam, and Peter Corke. A survey on terrain traversability analysis for autonomous ground ve- hicles: Methods...

  3. [10]

    Language models are few-shot learners

    Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-V oss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey ...

  4. [11]

    Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom

    Holger Caesar, Varun Bankiti, Alex H. Lang, Sourabh V ora, Venice Erin Liong, Qiang Xu, Anush Krishnan, Yu Pan, Giancarlo Baldan, and Oscar Beijbom. nuscenes: A multimodal dataset for autonomous driving. In CVPR, 2020

  5. [12]

    Evora: Deep evidential traversability learning for risk-aware off-road autonomy

    Xiaoyi Cai, Siddharth Ancha, Lakshay Sharma, Philip R Osteen, Bernadette Bucher, Stephen Phillips, Jiuguang Wang, Michael Everett, Nicholas Roy, and Jonathan P How. Evora: Deep evidential traversability learning for risk-aware off-road autonomy. IEEE Transactions on Robotics, 2024

  6. [13]

    How does it feel? self-supervised costmap learning for off-road vehicle traversability

    Mateo Guaman Castro, Samuel Triest, Wenshan Wang, Jason M Gregory, Felix Sanchez, John G Rogers, and Sebastian Scherer. How does it feel? self-supervised costmap learning for off-road vehicle traversability. In 2023 IEEE International Conference on Robotics and Automation (ICR...

  7. [14]

    Next Token Prediction Towards Multimodal Intelligence: A Comprehensive Survey, De- cember 2024

    Liang Chen, Zekun Wang, Shuhuai Ren, Lei Li, Haozhe Zhao, Yunshui Li, Zefan Cai, Hongcheng Guo, Lei Zhang, Yizhe Xiong, Yichi Zhang, Ruoyu Wu, Qingxiu Dong, Ge Zhang, Jian Yang, Lingwei Meng, Shujie Hu, Yulong Chen, Junyang Lin, Shuai Bai, Andreas Vlachos, Xu Tan, Minjia Zhang...

  8. [15]

    Decision Trans- former: Reinforcement Learning via Sequence Modeling

    Lili Chen, Kevin Lu, Aravind Rajeswaran, Kimin Lee, Aditya Grover, Michael Laskin, Pieter Abbeel, Ar- avind Srinivas, and Igor Mordatch. Decision Trans- former: Reinforcement Learning via Sequence Modeling. arXiv:2106.01345 [cs], June 2021

  9. [16]

    Generative pretraining from pixels

    Mark Chen, Alec Radford, Rewon Child, Jeffrey Wu, Heewoo Jun, David Luan, and Ilya Sutskever. Generative pretraining from pixels. In Hal Daum ´e III and Aarti Singh, editors, Proceedings of the 37th International Conference on Machine Learning , volume 119 of Pro- ceedings of ...

  10. [17]

    An Empirical Study of Training Self-Supervised Vision Transformers

    Xinlei Chen, Saining Xie, and Kaiming He. An Empirical Study of Training Self-Supervised Vision Transformers. In 2021 IEEE/CVF International Conference on Com- puter Vision (ICCV) , pages 9620–9629, Montreal, QC, Canada, October 2021. IEEE. ISBN 978-1-6654-2812-5. doi: 10.1109...

  11. [18]

    Hybrid imitative planning with geometric and predictive costs in off-road environ- ments

    Nitish Dashora, Daniel Shin, Dhruv Shah, Henry Leopold, David Fan, Ali Agha-Mohammadi, Nicholas Rhinehart, and Sergey Levine. Hybrid imitative planning with geometric and predictive costs in off-road environ- ments. In 2022 International Conference on Robotics and Automation (...

  12. [21]

    Learning to model and plan for wheeled mobility on vertically challenging terrain

    Aniket Datar, Chenhui Pan, and Xuesu Xiao. Learning to model and plan for wheeled mobility on vertically challenging terrain. IEEE Robotics and Automation Letters, 10(2):1505–1512, 2025

  13. [22]

    BERT: Pre-training of deep bidi- rectional transformers for language understanding

    Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: Pre-training of deep bidi- rectional transformers for language understanding. In Jill Burstein, Christy Doran, and Thamar Solorio, edi- tors, Proceedings of the 2019 Conference of the North American Chapte...

  14. [23]

    A note on two problems in connexion with graphs

    Edsger W Dijkstra. A note on two problems in connexion with graphs. Numerische mathematik , 1(1):269–271, 1959

  15. [24]

    Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation, August 2024

    Ria Doshi, Homer Walke, Oier Mees, Sudeep Dasari, and Sergey Levine. Scaling Cross-Embodied Learning: One Policy for Manipulation, Navigation, Locomotion and Aviation, August 2024

  16. [25]

    An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale, June 2021

    Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An Image is Worth 16x16 Words: Transformers for Image Recognition at...

  17. [26]

    VTNet: Visual Transformer Network for Object Goal Navigation, May 2021

    Heming Du, Xin Yu, and Liang Zheng. VTNet: Visual Transformer Network for Object Goal Navigation, May 2021

  18. [27]

    Tenenbaum, Leslie Kaelbling, Andy Zeng, and Jonathan Tompson

    Yilun Du, Mengjiao Yang, Pete Florence, Fei Xia, Ayzaan Wahid, Brian Ichter, Pierre Sermanet, Tianhe Yu, Pieter Abbeel, Joshua B. Tenenbaum, Leslie Kaelbling, Andy Zeng, and Jonathan Tompson. Video Language Planning, October 2023

  19. [28]

    Step: Stochastic traversability evaluation and planning for risk- aware off-road navigation

    David D Fan, Kyohei Otsu, Yuki Kubo, Anushri Dixit, Joel Burdick, and Ali-Akbar Agha-Mohammadi. Step: Stochastic traversability evaluation and planning for risk- aware off-road navigation. In Robotics: Science and Systems (RSS), 2021

  20. [29]

    Masked Autoencoders As Spatiotemporal Learners, May 2022

    Christoph Feichtenhofer, Haoqi Fan, Yanghao Li, and Kaiming He. Masked Autoencoders As Spatiotemporal Learners, May 2022

  21. [30]

    Foundation Models in Robotics: Applications, Chal- lenges, and the Future

    Roya Firoozi, Johnathan Tucker, Stephen Tian, Anirudha Majumdar, Jiankai Sun, Weiyu Liu, Yuke Zhu, Shuran Song, Ashish Kapoor, Karol Hausman, Brian Ichter, Danny Driess, Jiajun Wu, Cewu Lu, and Mac Schwa- ger. Foundation Models in Robotics: Applications, Chal- lenges, and the ...

  22. [31]

    How to train vision transformer on small-scale datasets? In 33rd British Machine Vision Conference Proceedings, BMVC 2022, 2022

    Hanan Gani, Muzammal Naseer, and Mohammad Yaqub. How to train vision transformer on small-scale datasets? In 33rd British Machine Vision Conference Proceedings, BMVC 2022, 2022

  23. [32]

    Multimodal Masked Autoencoders Learn Transferable Representations, May 2022

    Xinyang Geng, Hao Liu, Lisa Lee, Dale Schuurmans, Sergey Levine, and Pieter Abbeel. Multimodal Masked Autoencoders Learn Transferable Representations, May 2022

  24. [33]

    Zhang, Shaoqing Ren, and Jian Sun

    Kaiming He, X. Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pages 770–778, 2015. doi: 10.1109/cvpr.2016. 90

  25. [34]

    Masked autoencoders are scalable vision learners

    Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Doll´ar, and Ross Girshick. Masked autoencoders are scalable vision learners. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pages 16000–16009, 2022

  26. [35]

    Bridging nonlineari- ties and stochastic regularizers with gaussian error linear units, 2017

    Dan Hendrycks and Kevin Gimpel. Bridging nonlineari- ties and stochastic regularizers with gaussian error linear units, 2017

  27. [36]

    GAIA-1: A Generative World Model for Autonomous Driving, September 2023

    Anthony Hu, Lloyd Russell, Hudson Yeo, Zak Murez, George Fedoseev, Alex Kendall, Jamie Shotton, and Gianluca Corrado. GAIA-1: A Generative World Model for Autonomous Driving, September 2023

  28. [37]

    DrivingWorld: Constructing World Model for Au- tonomous Driving via Video GPT, December 2024

    Xiaotao Hu, Wei Yin, Mingkai Jia, Junyuan Deng, Xi- aoyang Guo, Qian Zhang, Xiaoxiao Long, and Ping Tan. DrivingWorld: Constructing World Model for Au- tonomous Driving via Video GPT, December 2024

  29. [38]

    Video Prediction Pol- icy: A Generalist Robot Policy with Predictive Visual Representations, December 2024

    Yucheng Hu, Yanjiang Guo, Pengchao Wang, Xiaoyu Chen, Yen-Jen Wang, Jianke Zhang, Koushil Sreenath, Chaochao Lu, and Jianyu Chen. Video Prediction Pol- icy: A Generalist Robot Policy with Predictive Visual Representations, December 2024

  30. [39]

    Goal-Guided Transformer-Enabled Reinforcement Learning for Efficient Autonomous Navigation

    Wenhui Huang, Yanxin Zhou, Xiangkun He, and Chen Lv. Goal-Guided Transformer-Enabled Reinforcement Learning for Efficient Autonomous Navigation. IEEE Transactions on Intelligent Transportation Systems , 25 (2):1832–1845, February 2024. ISSN 1558-0016. doi: 10.1109/TITS.2023.3312453

  31. [40]

    Rellis-3d dataset: Data, benchmarks and anal- ysis

    Peng Jiang, Philip Osteen, Maggie Wigness, and Srikanth Saripalli. Rellis-3d dataset: Data, benchmarks and anal- ysis. In 2021 IEEE international conference on robotics and automation (ICRA) , pages 1110–1116. IEEE, 2021

  32. [41]

    Johnson, Uday S

    Jacob J. Johnson, Uday S. Kalra, Ankit Bhatia, Linjun Li, Ahmed H. Qureshi, and Michael C. Yip. Motion Planning Transformers: A Motion Planning Framework for Mobile Robots, November 2022

  33. [42]

    Beyond Sight: Finetuning Generalist Robot Policies with Hetero- geneous Sensors via Language Grounding, January 2025

    Joshua Jones, Oier Mees, Carmelo Sferrazza, Kyle Sta- chowicz, Pieter Abbeel, and Sergey Levine. Beyond Sight: Finetuning Generalist Robot Policies with Hetero- geneous Sensors via Language Grounding, January 2025

  34. [43]

    VI-IKD: High-speed ac- curate off-road navigation using learned visual-inertial inverse kinodynamics

    Haresh Karnan, Kavan Singh Sikand, Pranav Atreya, Sadegh Rabiee, Xuesu Xiao, Garrett Warnell, Peter Stone, and Joydeep Biswas. VI-IKD: High-speed ac- curate off-road navigation using learned visual-inertial inverse kinodynamics. In 2022 IEEE/RSJ International Conference on Int...

  35. [44]

    DINO-Foresight Looking into the Future with DINO, December 2024

    Efstathios Karypidis, Ioannis Kakogeorgiou, Spyros Gi- daris, and Nikos Komodakis. DINO-Foresight Looking into the Future with DINO, December 2024

  36. [45]

    MarineFormer: A Spatio-Temporal Attention Model for USV Navigation in Dynamic Marine Environments, December 2024

    Ehsan Kazemi and Iman Soltani. MarineFormer: A Spatio-Temporal Attention Model for USV Navigation in Dynamic Marine Environments, December 2024

  37. [46]

    Daniel Lawson and Ahmed H. Qureshi. Control Transformer: Robot Navigation in Unknown Environ- ments Through PRM-Guided Return-Conditioned Se- quence Modeling. In 2023 IEEE/RSJ International Con- ference on Intelligent Robots and Systems (IROS) , pages 9324–9331, October 2023. ...

  38. [47]

    Learning terrain-aware kinodynamic model for autonomous off-road rally driving with model predictive path integral control

    Hojin Lee, Taekyung Kim, Jungwi Mun, and Wonsuk Lee. Learning terrain-aware kinodynamic model for autonomous off-road rally driving with model predictive path integral control. IEEE Robotics and Automation Letters, 2023

  39. [48]

    CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos, November 2024

    Xinhao Liu, Jintong Li, Yicheng Jiang, Niranjan Su- jay, Zhicheng Yang, Juexiao Zhang, John Abanes, Jing Zhang, and Chen Feng. CityWalker: Learning Embodied Urban Navigation from Web-Scale Videos, November 2024

  40. [49]

    Efficient training of visual transformers with small datasets

    Yahui Liu, Enver Sangineto, Wei Bi, Nicu Sebe, Bruno Lepri, and Marco Nadai. Efficient training of visual transformers with small datasets. Advances in Neural Information Processing Systems, 34:23818–23830, 2021

  41. [50]

    Decoupled Weight Decay Regularization, January 2019

    Ilya Loshchilov and Frank Hutter. Decoupled Weight Decay Regularization, January 2019

  42. [51]

    nGPT: Normalized Transformer with Representation Learning on the Hypersphere, October 2024

    Ilya Loshchilov, Cheng-Ping Hsieh, Simeng Sun, and Boris Ginsburg. nGPT: Normalized Transformer with Representation Learning on the Hypersphere, October 2024

  43. [52]

    Discrete Rep- resentations Strengthen Vision Transformer Robustness, April 2022

    Chengzhi Mao, Lu Jiang, Mostafa Dehghani, Carl V on- drick, Rahul Sukthankar, and Irfan Essa. Discrete Rep- resentations Strengthen Vision Transformer Robustness, April 2022

  44. [53]

    GPT-Driver: Learning to Drive with GPT, October 2023

    Jiageng Mao, Yuxi Qian, Hang Zhao, and Yue Wang. GPT-Driver: Learning to Drive with GPT, October 2023

  45. [54]

    Wild Visual Navigation: Fast Traversability Learning via Pre-Trained Models and On- line Self-Supervision, April 2024

    Mat ´ıas Mattamala, Jonas Frey, Piotr Libera, Nived Che- brolu, Georg Martius, Cesar Cadena, Marco Hutter, and Maurice Fallon. Wild Visual Navigation: Fast Traversability Learning via Pre-Trained Models and On- line Self-Supervision, April 2024

  46. [55]

    Elevation mapping for locomotion and navigation using gpu

    Takahiro Miki, Lorenz Wellhausen, Ruben Grandia, Fabian Jenelten, Timon Homberger, and Marco Hutter. Elevation mapping for locomotion and navigation using gpu. In 2022 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 2273–2280. IEEE, 2022

  47. [56]

    VertiEncoder: Self-Supervised Kinodynamic Representation Learning on Vertically Challenging Terrain, September 2024

    Mohammad Nazeri, Aniket Datar, Anuj Pokhrel, Chenhui Pan, Garrett Warnell, and Xuesu Xiao. VertiEncoder: Self-Supervised Kinodynamic Representation Learning on Vertically Challenging Terrain, September 2024

  48. [57]

    V ANP: Learning Where to See for Navigation with Self-Supervised Vision-Action Pre- Training

    Mohammad Nazeri, Junzhe Wang, Amirreza Payandeh, and Xuesu Xiao. V ANP: Learning Where to See for Navigation with Self-Supervised Vision-Action Pre- Training. In 2024 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 2741–2746, Abu Dhabi, United...

  49. [58]

    Ex- ploring Reflective Limitation of Behavior Cloning in Autonomous Vehicles

    Mohammad Hossein Nazeri and Mahdi Bohlouli. Ex- ploring Reflective Limitation of Behavior Cloning in Autonomous Vehicles. In 2021 IEEE International Conference on Data Mining (ICDM) , pages 1252–1257, December 2021. doi: 10.1109/ICDM51629.2021.00153

  50. [59]

    Octo: An Open-Source Generalist Robot Policy, May 2024

    Octo Model Team, Dibya Ghosh, Homer Walke, Karl Pertsch, Kevin Black, Oier Mees, Sudeep Dasari, Joey Hejna, Tobias Kreiman, Charles Xu, Jianlan Luo, You Liang Tan, Pannag Sanketi, Quan Vuong, Ted Xiao, Dorsa Sadigh, Chelsea Finn, and Sergey Levine. Octo: An Open-Source General...

  51. [60]

    DINOv2: Learning Robust Visual Features without Supervision, April 2023

    Maxime Oquab, Timoth ´ee Darcet, Th ´eo Moutakanni, Huy V o, Marc Szafraniec, Vasil Khalidov, Pierre Fer- nandez, Daniel Haziza, Francisco Massa, Alaaeldin El- Nouby, Mahmoud Assran, Nicolas Ballas, Wojciech Galuba, Russell Howes, Po-Yao Huang, Shang-Wen Li, Ishan Misra, Micha...

  52. [61]

    Fast local planning and mapping in unknown off-road terrain

    Timothy Overbye and Srikanth Saripalli. Fast local planning and mapping in unknown off-road terrain. In 2020 IEEE International Conference on Robotics and Automation (ICRA), pages 5912–5918. IEEE, 2020

  53. [62]

    Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration

    Abby O’Neill, Abdul Rehman, Abhiram Maddukuri, Ab- hishek Gupta, Abhishek Padalkar, Abraham Lee, Acorn Pooley, Agrim Gupta, Ajay Mandlekar, Ajinkya Jain, et al. Open x-embodiment: Robotic learning datasets and rt-x models: Open x-embodiment collaboration. In 2024 IEEE Internat...

  54. [63]

    Tra- verse the Non-Traversable: Estimating Traversability for Wheeled Mobility on Vertically Challenging Terrain, September 2024

    Chenhui Pan, Aniket Datar, Anuj Pokhrel, Matthew Choulas, Mohammad Nazeri, and Xuesu Xiao. Tra- verse the Non-Traversable: Estimating Traversability for Wheeled Mobility on Vertically Challenging Terrain, September 2024

  55. [64]

    Imitation learning for agile autonomous driving

    Yunpeng Pan, Ching-An Cheng, Kamil Saigol, Keuntaek Lee, Xinyan Yan, Evangelos A Theodorou, and Byron Boots. Imitation learning for agile autonomous driving. The International Journal of Robotics Research , 2020

  56. [65]

    Viorica P ˘atr˘aucean, Xu Owen He, Joseph Heyward, Chuhan Zhang, Mehdi S. M. Sajjadi, George-Cristian Muraru, Artem Zholus, Mahdi Karami, Ross Goroshin, Yutian Chen, Simon Osindero, Jo˜ao Carreira, and Razvan Pascanu. TRecViT: A Recurrent Video Transformer, December 2024

  57. [66]

    Transformers for Image-Goal Naviga- tion, May 2024

    Nikhilanj Pelluri. Transformers for Image-Goal Naviga- tion, May 2024

  58. [67]

    FAST: Efficient Action To- kenization for Vision-Language-Action Models, January 2025

    Karl Pertsch, Kyle Stachowicz, Brian Ichter, Danny Driess, Suraj Nair, Quan Vuong, Oier Mees, Chelsea Finn, and Sergey Levine. FAST: Efficient Action To- kenization for Vision-Language-Action Models, January 2025

  59. [68]

    CAHSOR: Competence-aware high-speed off-road ground navigation in SE (3)

    Anuj Pokhrel, Aniket Datar, Mohammad Nazeri, and Xuesu Xiao. CAHSOR: Competence-aware high-speed off-road ground navigation in SE (3). IEEE Robotics and Automation Letters, 9(11):9653–9660, 2024

  60. [69]

    Pomerleau

    Dean A. Pomerleau. ALVINN: An autonomous land vehicle in a neural network. In D. Touretzky, editor, Advances in Neural Information Processing Systems , volume 1. Morgan-Kaufmann, 1988

  61. [70]

    Improving language understanding by generative pre-training

    Alec Radford, Karthik Narasimhan, Tim Salimans, and Ilya Sutskever. Improving language understanding by generative pre-training. 2018

  62. [71]

    Language models are unsupervised multitask learners

    Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. Language models are unsupervised multitask learners. 2019

  63. [72]

    An Empirical Study of Autoregressive Pre-training from Videos, January 2025

    Jathushan Rajasegaran, Ilija Radosavovic, Rahul Ravis- hankar, Yossi Gandelsman, Christoph Feichtenhofer, and Jitendra Malik. An Empirical Study of Autoregressive Pre-training from Videos, January 2025

  64. [73]

    Neural Machine Translation of Rare Words with Sub- word Units, June 2016

    Rico Sennrich, Barry Haddow, and Alexandra Birch. Neural Machine Translation of Rare Words with Sub- word Units, June 2016

  65. [74]

    Learning Off-Road Terrain Traversability with Self-Supervisions Only, May 2023

    Junwon Seo, Sungdae Sim, and Inwook Shim. Learning Off-Road Terrain Traversability with Self-Supervisions Only, May 2023

  66. [75]

    Masked World Models for Visual Control, May 2023

    Younggyo Seo, Danijar Hafner, Hao Liu, Fangchen Liu, Stephen James, Kimin Lee, and Pieter Abbeel. Masked World Models for Visual Control, May 2023

  67. [76]

    Multi-View Masked World Models for Visual Robotic Manipulation, February 2023

    Younggyo Seo, Junsu Kim, Stephen James, Kimin Lee, Jinwoo Shin, and Pieter Abbeel. Multi-View Masked World Models for Visual Robotic Manipulation, February 2023

  68. [77]

    Ramp: A risk- aware mapping and planning pipeline for fast off-road ground robot navigation

    Lakshay Sharma, Michael Everett, Donggun Lee, Xiaoyi Cai, Philip Osteen, and Jonathan P How. Ramp: A risk- aware mapping and planning pipeline for fast off-road ground robot navigation. In 2023 IEEE International Conference on Robotics and Automation (ICRA) , pages 5730–5736. ...

  69. [78]

    Robot adaptation to unstructured terrains by joint representation and apprenticeship learning

    Sriram Siva, Maggie Wigness, John Rogers, and Hao Zhang. Robot adaptation to unstructured terrains by joint representation and apprenticeship learning. In Robotics: Science and Systems (RSS) , 2019

  70. [79]

    Nauts: Negotiation for adapta- tion to unstructured terrain surfaces

    Sriram Siva, Maggie Wigness, John G Rogers, Long Quang, and Hao Zhang. Nauts: Negotiation for adapta- tion to unstructured terrain surfaces. In 2022 IEEE/RSJ International Conference on Intelligent Robots and Sys- tems (IROS), pages 1733–1740. IEEE, 2022

  71. [80]

    Improving off-road planning techniques with learned costs from physical interactions

    Matthew Sivaprakasam, Samuel Triest, Wenshan Wang, Peng Yin, and Sebastian Scherer. Improving off-road planning techniques with learned costs from physical interactions. In 2021 IEEE International Conference on Robotics and Automation (ICRA) , pages 4844–4850. IEEE, 2021

  72. [81]

    How to train your ViT? Data, augmentation, and regular- ization in vision transformers

    Andreas Peter Steiner, Alexander Kolesnikov, Xiaohua Zhai, Ross Wightman, Jakob Uszkoreit, and Lucas Beyer. How to train your ViT? Data, augmentation, and regular- ization in vision transformers. Transactions on Machine Learning Research, 2022. ISSN 2835-8856

  73. [82]

    Scalabil- ity in perception for autonomous driving: Waymo open dataset

    Pei Sun, Henrik Kretzschmar, Xerxes Dotiwalla, Aure- lien Chouard, Vijaysai Patnaik, Paul Tsui, James Guo, Yin Zhou, Yuning Chai, Benjamin Caine, et al. Scalabil- ity in perception for autonomous driving: Waymo open dataset. In Proceedings of the IEEE/CVF conference on compute...

  74. [83]

    Tartandrive: A large-scale dataset for learning off-road dynamics models

    Samuel Triest, Matthew Sivaprakasam, Sean J Wang, Wenshan Wang, Aaron M Johnson, and Sebastian Scherer. Tartandrive: A large-scale dataset for learning off-road dynamics models. In 2022 International Confer- ence on Robotics and Automation (ICRA) , pages 2546–

  75. [84]

    Attention is all you need

    Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. Advances in neural information processing systems , 30, 2017

  76. [85]

    Nav- Former: A Transformer Architecture for Robot Target- Driven Navigation in Unknown and Dynamic Environ- ments

    Haitong Wang, Aaron Hao Tan, and Goldie Nejat. Nav- Former: A Transformer Architecture for Robot Target- Driven Navigation in Unknown and Dynamic Environ- ments. 2024. doi: 10.48550/ARXIV .2402.06838

  77. [86]

    Navigating the Landscape of Large Lan- guage Models: A Comprehensive Review and Analysis of Paradigms and Fine-Tuning Strategies, April 2024

    Benjue Weng. Navigating the Landscape of Large Lan- guage Models: A Comprehensive Review and Analysis of Paradigms and Fine-Tuning Strategies, April 2024

  78. [87]

    A rugd dataset for autonomous navigation and visual perception in unstructured outdoor environments

    Maggie Wigness, Sungmin Eum, John G Rogers, David Han, and Heesung Kwon. A rugd dataset for autonomous navigation and visual perception in unstructured outdoor environments. In 2019 IEEE/RSJ International Confer- ence on Intelligent Robots and Systems (IROS) , pages 5000–5007....

  79. [88]

    Model predictive path integral control: From theory to parallel computation

    Grady Williams, Andrew Aldrich, and Evangelos A Theodorou. Model predictive path integral control: From theory to parallel computation. Journal of Guidance, Control, and Dynamics , 2017

  80. [89]

    Dolan, and Guanya Shi

    Wenli Xiao, Haoru Xue, Tony Tao, Dvij Kalaria, John M. Dolan, and Guanya Shi. AnyCar to Anywhere: Learn- ing Universal Dynamics Model for Agile and Adaptive Mobility, September 2024

  81. [90]

    Learning inverse kinodynamics for accurate high-speed off-road navigation on unstructured terrain

    Xuesu Xiao, Joydeep Biswas, and Peter Stone. Learning inverse kinodynamics for accurate high-speed off-road navigation on unstructured terrain. IEEE Robotics and Automation Letters, 6(3):6054–6060, 2021

  82. [91]

    Motion planning and control for mobile robot navigation using machine learning: A survey

    Xuesu Xiao, Bo Liu, Garrett Warnell, and Peter Stone. Motion planning and control for mobile robot navigation using machine learning: A survey. Autonomous Robots, 46(5):569–597, June 2022. ISSN 0929-5593, 1573-7527. doi: 10.1007/s10514-022-10039-8

  83. [92]

    On layer normalization in the transformer architecture

    Ruibin Xiong, Yunchang Yang, Di He, Kai Zheng, Shuxin Zheng, Chen Xing, Huishuai Zhang, Yanyan Lan, Liwei Wang, and Tieyan Liu. On layer normalization in the transformer architecture. In International Conference on Machine Learning, pages 10524–10533. PMLR, 2020. doi: 10.5555/...

  84. [93]

    Prince, and Yanshuai Cao

    Peng Xu, Dhruv Kumar, Wei Yang, Wenjie Zi, Keyi Tang, Chenyang Huang, Jackie Chi Kit Cheung, Si- mon J.D. Prince, and Yanshuai Cao. Optimizing Deeper Transformers on Small Datasets. In Proceedings of the 59th Annual Meeting of the Association for Com- putational Linguistics an...

  85. [94]

    Scaling vision transformers

    Xiaohua Zhai, Alexander Kolesnikov, Neil Houlsby, and Lucas Beyer. Scaling vision transformers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) , pages 12104–12113, June 2022

  86. [95]

    Root mean square layer normalization

    Biao Zhang and Rico Sennrich. Root mean square layer normalization. In Advances in Neural Information Processing Systems 32 , Vancouver, Canada, 2019

  87. [96]

    NaviFormer: A Data- Driven Robot Navigation Approach via Sequence Mod- eling and Path Planning with Safety Verification

    Xuyang Zhang, Ziyang Feng, Quecheng Qiu, Yu’an Chen, Bei Hua, and Jianmin Ji. NaviFormer: A Data- Driven Robot Navigation Approach via Sequence Mod- eling and Path Planning with Safety Verification. In 2024 IEEE International Conference on Robotics and Automation (ICRA), pages...

  88. [97]

    DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning, November 2024

    Gaoyue Zhou, Hengkai Pan, Yann LeCun, and Lerrel Pinto. DINO-WM: World Models on Pre-trained Visual Features enable Zero-shot Planning, November 2024. APPENDIX A MODEL ARCHITECTURE TABLE II: V ERTI FORMER Architecture Parameters. VERTI ENCODER Layers 6 Normalization RMSNorm [9...

  89. [1703]

    PMLR, 2020-07-13/2020-07-18

Pith tools

Reviewed August 9, 2026 · model on record in the stance chip above.