Pith. sign in

REVIEW 2 major objections 56 references

GUIDE: Goal-Initialized Directional Understanding for End-to-End Visual Navigation

T0 review · 2 major / 0 minor · reviewed 2026-06-27 · grok-4.3

Pith's one-line read A goal given once at the start is enough for an end-to-end policy to navigate dense clutter and mazes on a quadruped.

desk verdict GUIDE frames a goal-initialized navigation problem for legged robots using proprioceptive spatial anchors, but the abstract supplies zero numbers or comparisons so the claims stay uncheckable. read the letter →

arxiv 2606.10832 v1 pith:DV4ZUE46 submitted 2026-06-09 cs.RO

classification cs.RO
keywords end-to-endvisualnavigationreinforcementlearningleggedrobotsegomotionestimationspatialmemorygoal-initializedquadrupedlocomotionproprioceptivehistory
verification ladder T0 review T1 audit T2 compute T3 formal

The pith

A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.

The reading

The paper tries to establish that legged-robot visual navigation can succeed without continuous goal updates from external state estimators. GUIDE builds an internal directional sense by feeding multi-frequency proprioceptive history into a spatial anchor predictor and raw depth into a local geometry encoder. A reader would care because this removes extra sensors and computation while avoiding myopic traps in partial views. The claim is tested in both simulation and hardware on a quadruped moving through clutter and structured layouts. If correct, robots could operate with only a single initial target and intrinsic memory.

What carries the argument

The spatial anchor predictor, which turns multi-frequency proprioceptive history into egomotion representations that sustain long-horizon spatial context under partial observability.

What would settle it

A controlled trial in which the robot receives only single-frequency proprioception and then repeatedly loses its way or enters dead-end loops in the same maze layouts would falsify the claim.

Watch

Extended reading notes

Core claim

GUIDE is a fully end-to-end reinforcement learning framework that cultivates internal directional awareness by incorporating a spatial anchor predictor leveraging multi-frequency proprioceptive history to extract egomotion representations, thereby maintaining a persistent long-horizon spatial context for navigation, while simultaneously using raw depth streams to perceive local environmental geometry.

Load-bearing premise

Multi-frequency proprioceptive history by itself is enough for the spatial anchor predictor to produce egomotion signals that keep directional awareness intact over long episodes.

Editorial extensions

If this is right

  • The deployed policy navigates dense clutter and structured mazes without any further goal signals or prior maps.
  • Reliable egomotion and directional awareness emerge from intrinsic spatial memory alone.
  • The same network runs end-to-end in both simulation and on real quadruped hardware.
  • No hierarchical state estimation module is required after the initial goal is supplied.

Reading between the lines

Editorial extensions of the paper, not claims the author makes directly.

  • The method may lower the sensor budget for field-deployed legged robots by removing external pose estimators.
  • Similar internal-memory designs could be tested on other platforms whose proprioception shares comparable frequency content.
  • Longer autonomous missions become feasible if the spatial anchor continues to function when depth is intermittently lost.
  • The approach directly addresses partial-observability memory problems that appear in many other sequential decision tasks.
Share X Bluesky LinkedIn Reddit HN

Editorial analysis

A structured set of objections, weighed in public.

Desk editor's note, referee report, simulated authors' rebuttal, and a circularity audit.

Referee Report

2 major / 0 minor

Summary. The manuscript proposes GUIDE, a fully end-to-end reinforcement learning framework for goal-initialized visual navigation on legged robots. In this setting the goal is supplied only at episode start; the policy must then rely on intrinsic spatial memory. GUIDE combines a spatial anchor predictor that processes multi-frequency proprioceptive history to extract egomotion representations and maintain long-horizon directional context with raw depth inputs for local geometry. The authors claim that the resulting policy navigates dense clutter and structured mazes without subsequent goal updates or prior maps, and they report evaluation in both simulation and real-world quadruped experiments.

Significance. If the empirical claims are substantiated, the work would demonstrate that proprioceptive history alone can furnish a persistent spatial anchor sufficient for long-horizon navigation, thereby removing the need for continuous external state-estimation modules. This could simplify deployment pipelines and reduce sensory/computational overhead on resource-limited platforms. The absence of any quantitative results in the supplied text, however, prevents assessment of whether the claimed reliability is actually achieved.

major comments (2)
  1. [Abstract] Abstract: the central claim that 'experiments show that GUIDE learns reliable egomotion and directional awareness, enabling a fully end-to-end deployed policy to safely navigate through dense clutter and structured mazes' is unsupported by any quantitative metrics, success rates, baselines, ablation results, or statistical comparisons. Without these data the soundness of the contribution cannot be evaluated.
  2. [Abstract] The weakest assumption identified—that multi-frequency proprioceptive history alone suffices for the spatial anchor predictor to sustain long-horizon context under partial observability—is stated but never tested against alternative proprioceptive encodings or against ground-truth egomotion; no ablation or sensitivity analysis is supplied to support this design choice.

Simulated Author's Rebuttal

2 responses · 0 unresolved

We thank the referee for their detailed review and constructive comments on our manuscript. We acknowledge that the abstract claims require stronger quantitative backing in the provided text and that additional ablations would better support the design choices. We will revise the manuscript to address these points directly.

read point-by-point responses
  1. Referee: [Abstract] Abstract: the central claim that 'experiments show that GUIDE learns reliable egomotion and directional awareness, enabling a fully end-to-end deployed policy to safely navigate through dense clutter and structured mazes' is unsupported by any quantitative metrics, success rates, baselines, ablation results, or statistical comparisons. Without these data the soundness of the contribution cannot be evaluated.

    Authors: We agree that the abstract's claims must be supported by quantitative evidence for the contribution to be properly evaluated. The full manuscript contains simulation and real-world results with success rates, baseline comparisons, and statistical details; however, since these were not evident in the supplied text, we will revise the abstract to incorporate key metrics (e.g., navigation success rates in clutter and mazes) and ensure the results section is clearly linked to the claims. revision: yes

  2. Referee: [Abstract] The weakest assumption identified—that multi-frequency proprioceptive history alone suffices for the spatial anchor predictor to sustain long-horizon context under partial observability—is stated but never tested against alternative proprioceptive encodings or against ground-truth egomotion; no ablation or sensitivity analysis is supplied to support this design choice.

    Authors: We agree that explicit ablations and sensitivity analyses would strengthen the justification for the multi-frequency proprioceptive encoding. We will add these experiments in the revised manuscript, including comparisons to alternative encodings (e.g., single-frequency or raw IMU) and validation against ground-truth egomotion where available, to directly test the assumption under partial observability. revision: yes

Circularity Check

0 steps flagged · score 0.0 of 10

No significant circularity

full rationale

The manuscript is a high-level architectural proposal for an RL navigation policy. No equations, parameter-fitting procedures, uniqueness theorems, or derivation chains appear in the abstract or descriptive sections. Claims about the spatial anchor predictor are presented as empirical outcomes of training rather than reductions to prior fitted quantities or self-citations. The work is therefore self-contained against external benchmarks with no detectable circular steps.

Assumptions & free parameters 0 free parameters · 0 assumptions · 0 invented entities

The abstract contains no explicit free parameters, axioms, or invented entities; the framework is described conceptually without numerical fitting or new postulated objects.

how reviews work

0 comments
Cite this review

Pith. "Pith review of GUIDE: Goal-Initialized Directional Understanding for End-to-End Visual Navigation." pith.science (2026). https://pith.science/paper/DV4ZUE46

@misc{pith2026260610832,
  author       = {Pith},
  title        = {Pith review of: GUIDE: Goal-Initialized Directional Understanding for End-to-End Visual Navigation},
  year         = {2026},
  howpublished = {\url{https://pith.science/paper/DV4ZUE46}},
  note         = {Machine review of arXiv:2606.10832}
}
read the original abstract

Learning-based visual navigation for legged robots typically relies on continuous goal updates from hierarchical state estimation to provide a persistent directional reference. This reliance incurs additional sensory and computational overhead and deviates from fully end-to-end mobile autonomy. Furthermore, under partial observability, policies are prone to learn myopic behaviors, easily becoming trapped in dead ends and complex structural layouts. To address these limitations, we investigate a goal-initialized navigation setting, where the target is provided only once at the beginning of an episode, requiring the robot to operate based on intrinsic spatial memory without subsequent goal updates from external modules. In this work, we propose GUIDE, a fully end-to-end reinforcement learning framework designed to cultivate internal directional awareness. Specifically, GUIDE incorporates a spatial anchor predictor that leverages multi-frequency proprioceptive history to extract egomotion representations, thereby maintaining a persistent long-horizon spatial context for navigation. Concurrently, it utilizes raw depth streams to perceive local environmental geometry. We evaluate the proposed framework across both simulation and real-world scenarios on a quadruped robot. Experiments show that GUIDE learns reliable egomotion and directional awareness, enabling a fully end-to-end deployed policy to safely navigate through dense clutter and structured mazes without subsequent goal guidance or prior maps.

Figures

Figures reproduced from arXiv: 2606.10832 by the authors.

Figure 1
Figure 1. Our GUIDE framework cultivates internal egomotion and directional awareness for end [PITH_FULL_IMAGE:figures/full_fig_p001_1.png] view at source ↗
Figure 2
Figure 2. Overview of the GUIDE framework. Multi-frequency proprioceptive history is pro￾cessed into proprioceptive tokens, which are supervised to predict spatial anchor vectors to cultivate egomotion and directional awareness. Concurrently, depth buffers are encoded and fused with these tokens via cross-attention to yield spatial latents. Finally, the Actor aggregates these representations with the latest proprioceptive sta… view at source ↗
Figure 3
Figure 3. (a1, b1): Cluttered and maze terrains, with curves in gradient color indicating the robot’s trajectories. (a2, b2): Corresponding image inputs for the privileged critic, displaying the two￾channel downsampled maps (Mocc and Mexp) within the same image. 3.3 Reward Design and Multi-Critic Value Estimation Reward Function Design Our navigation reward design combines sparse task signals with dense progress guidance. The… view at source ↗
Figures from the paper (7 more)
Figure 4
Figure 4. Figure 4: Quantitative results of real-world en￾vironments. The robot will halt when the inter￾nally estimated distance to the goal is lower than 0.2m. GUIDE achieves the best results across all metrics. Besides, an average FD around 0.2 and 0.3m represents a 0-to-0.1 meter aver…
Figure 5
Figure 5. Figure 5: Real-world deployments. GUIDE successfully navigates through (a) in-lab cluttered environments and (b) mazes, as well as unstructured scenarios like (c) long office corridors and outdoor grasslands with (d) dense vegetation and (e) dynamic obstacles. Top-left insets di…
Figure 6
Figure 6. Figure 6: A representative failure case under lim [PITH_FULL_IMAGE:figures/full_fig_p008_6.png]
Figure 7
Figure 7. Figure 7: Depth observation processing pipeline. (a) Raw depth image rendered by NVIDIA Warp. (b) Depth image after additive Gaussian noise. (c) Depth Dropout. (d) Quantization. (e) Final policy input after clipping, normalization, cropping, and average-pooling downsampling. Vis…
Figure 8
Figure 8. Figure 8: Training curves for the critic abla￾tion. Compared to the w/o Multi-Critic vari￾ant, GUIDE demonstrates accelerated conver￾gence and achieves a higher total episode reward, highlighting the effectiveness of the Multi-Critic (MuC) architecture. Critic Ablation Analysis.…
Figure 9
Figure 9. Figure 9: Terrain curriculum progression. The terrain difficulty progressively increases from left to right for both cluttered and maze environments. These layouts also serve as examples of the eval￾uation benchmarks introduced in Section 4.1. Training Curriculum. Learning to na…
Figure 10
Figure 10. Figure 10: Additional real-world maze deployments. The figure illustrates three distinct spawn￾to-goal navigation trials within a 12 m×12 m physical maze. For each trial (rows a–c), the left panels (a1, b1, c1) display the wide view of the environment, while the right panels (a2…

Discussion (0). Continue with ORCID to comment.

Reference graph

Works this paper leans on

56 extracted references · 6 canonical work pages

  1. [1]

    Hoeller, N

    D. Hoeller, N. Rudin, D. Sako, and M. Hutter. Anymal parkour: Learning agile navigation for quadrupedal robots.Science Robotics, 9(88):eadi7566, 2024

  2. [2]

    Cheng, K

    X. Cheng, K. Shi, A. Agarwal, and D. Pathak. Extreme parkour with legged robots. In2024 IEEE International Conference on Robotics and Automation (ICRA), pages 11443–11450. IEEE, 2024

  3. [3]

    Y . Wu, J. Kuang, S. Khorshidi, X. Niu, L. Klingbeil, M. Bennewitz, and H. Kuhlmann. Doglegs: Robust proprioceptive state estimation for legged robots using multiple leg-mounted imus. In2025 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 8856–8863, 2025

  4. [4]

    F. Yang, P. Frivik, D. Hoeller, C. Wang, C. Cadena, and M. Hutter. Spatially-enhanced recur- rent memory for long-range mapless navigation via end-to-end reinforcement learning.The International Journal of Robotics Research, page 02783649251401926, 2025

  5. [5]

    HiPAN: Hierarchical Posture-Adaptive Navigation for Quadruped Robots in Unstructured 3D Environments

    J. Jeong, M. Yoon, S. Choi, H. Shin, T. Yang, and S.-e. Yoon. Hipan: Hierarchical posture- adaptive navigation for quadruped robots in unstructured 3d environments.arXiv preprint arXiv:2604.26504, 2026

  6. [6]

    Vernaza, M

    P. Vernaza, M. Likhachev, S. Bhattacharya, S. Chitta, A. Kushleyev, and D. D. Lee. Search- based planning for a legged robot over rough terrain. In2009 IEEE International Conference on Robotics and Automation, pages 2380–2387. IEEE, 2009

  7. [7]

    B. Yang, L. Wellhausen, T. Miki, M. Liu, and M. Hutter. Real-time optimal navigation plan- ning using learned motion costs. In2021 IEEE International Conference on Robotics and Automation (ICRA), pages 9283–9289. IEEE, 2021

  8. [8]

    Gaertner, M

    M. Gaertner, M. Bjelonic, F. Farshidian, and M. Hutter. Collision-free mpc for legged robots in static and dynamic scenes. In2021 IEEE International Conference on Robotics and Au- tomation (ICRA), pages 8266–8272. IEEE, 2021

Show all 56 references
  1. [9]

    C. Cao, H. Zhu, F. Yang, Y . Xia, H. Choset, J. Oh, and J. Zhang. Autonomous exploration development environment and the planning algorithms. In2022 International Conference on Robotics and Automation (ICRA), pages 8921–8928. IEEE, 2022

  2. [10]

    Jiang, H

    P. Jiang, H. Liu, J. Jin, W. Wang, and X. Li. Learning to evolve: Multi-modal interactive fields for robust humanoid navigation in dynamic environments. InRobotics: Science and Systems, 2026

  3. [11]

    Pfeiffer, M

    M. Pfeiffer, M. Schaeuble, J. I. Nieto, R. Y . Siegwart, and C. Cadena. From perception to de- cision: A data-driven approach to end-to-end motion planning for autonomous ground robots. 2017 IEEE International Conference on Robotics and Automation (ICRA), pages 1527–1533, 2016

  4. [12]

    Loquercio, E

    A. Loquercio, E. Kaufmann, R. Ranftl, M. M ¨uller, V . Koltun, and D. Scaramuzza. Learning high-speed flight in the wild.Science Robotics, 6(59):eabg5810, 2021

  5. [13]

    L. Tai, G. Paolo, and M. Liu. Virtual-to-real deep reinforcement learning: Continuous con- trol of mobile robots for mapless navigation. In2017 IEEE/RSJ international conference on intelligent robots and systems (IROS), pages 31–36. IEEE, 2017. 9

  6. [14]

    Wijmans, A

    E. Wijmans, A. Kadian, A. S. Morcos, S. Lee, I. Essa, D. Parikh, M. Savva, and D. Batra. Dd-ppo: Learning near-perfect pointgoal navigators from 2.5 billion frames. InInternational Conference on Learning Representations, 2019

  7. [15]

    Zhu and M

    W. Zhu and M. Hayashibe. A hierarchical deep reinforcement learning framework with high efficiency and generalization for fast and safe navigation.IEEE Transactions on industrial Electronics, 70(5):4962–4971, 2022

  8. [16]

    Y . Zhu, R. Mottaghi, E. Kolve, J. J. Lim, A. Gupta, L. Fei-Fei, and A. Farhadi. Target-driven visual navigation in indoor scenes using deep reinforcement learning. In2017 IEEE interna- tional conference on robotics and automation (ICRA), pages 3357–3364. IEEE, 2017

  9. [17]

    H. Shi, L. Shi, M. Xu, and K.-S. Hwang. End-to-end navigation strategy with deep rein- forcement learning for mobile robots.IEEE Transactions on Industrial Informatics, 16(4): 2393–2402, 2019

  10. [18]

    Huang, Y

    W. Huang, Y . Zhou, X. He, and C. Lv. Goal-guided transformer-enabled reinforcement learn- ing for efficient autonomous navigation.IEEE Transactions on Intelligent Transportation Sys- tems, 25(2):1832–1845, 2023

  11. [19]

    W ¨ohlke, F

    J. W ¨ohlke, F. Schmitt, and H. van Hoof. Hierarchies of planning and reinforcement learning for robot navigation. In2021 IEEE international conference on robotics and automation (ICRA), pages 10682–10688. IEEE, 2021

  12. [20]

    Y . Gao, J. Wu, X. Yang, and Z. Ji. Efficient hierarchical reinforcement learning for mapless navigation with predictive neighbouring space scoring.IEEE Transactions on Automation Science and Engineering, 21(4):5457–5472, 2023

  13. [21]

    J. Gao, X. Pang, Q. Liu, and Y . Li. Hierarchical reinforcement learning for safe mapless navigation with congestion estimation. In2025 IEEE International Conference on Robotics and Automation (ICRA), pages 8849–8855. IEEE, 2025

  14. [22]

    Pfeiffer, S

    M. Pfeiffer, S. Shukla, M. Turchetta, C. Cadena, A. Krause, R. Siegwart, and J. Nieto. Re- inforced imitation: Sample efficient deep reinforcement learning for mapless navigation by leveraging prior demonstrations.IEEE Robotics and Automation Letters, 3(4):4423–4430, 2018

  15. [23]

    J. Ye, D. Batra, A. Das, and E. Wijmans. Auxiliary tasks and exploration enable objectgoal navigation. InProceedings of the IEEE/CVF international conference on computer vision, pages 16117–16126, 2021

  16. [24]

    J. Lee, M. Bjelonic, A. Reske, L. Wellhausen, T. Miki, and M. Hutter. Learning robust au- tonomous navigation and locomotion for wheeled-legged robots.Science Robotics, 9(89): eadi9641, 2024

  17. [25]

    Q. Yuan, Z. Cao, M. Cao, and K. Li. Reasan: Learning reactive safe navigation for legged robots.arXiv preprint arXiv:2512.09537, 2025

  18. [26]

    Sadat, S

    A. Sadat, S. Casas, M. Ren, X. Wu, P. Dhawan, and R. Urtasun. Perceive, predict, and plan: Safe motion planning through interpretable semantic representations. InEuropean Conference on Computer Vision, pages 414–430. Springer, 2020

  19. [27]

    D. Shah, A. Sridhar, A. Bhorkar, N. Hirose, and S. Levine. Gnm: A general navigation model to drive any robot. In2023 IEEE International Conference on Robotics and Automation (ICRA), pages 7226–7233. IEEE, 2023

  20. [28]

    D. Shah, A. K. Sridhar, N. Dashora, K. Stachowicz, K. Black, N. Hirose, and S. Levine. Vint: A foundation model for visual navigation. InConference on Robot Learning, 2023. 10

  21. [29]

    F. Yang, C. Wang, C. Cadena, and M. Hutter. iplanner: Imperative path planning. InRobotics: Science and Systems, 2023

  22. [30]

    P. Roth, J. Nubert, F. Yang, M. Mittal, and M. Hutter. Viplanner: Visual semantic impera- tive learning for local navigation. In2024 IEEE International Conference on Robotics and Automation (ICRA), pages 5243–5249. IEEE, 2024

  23. [31]

    Hoeller, L

    D. Hoeller, L. Wellhausen, F. Farshidian, and M. Hutter. Learning a state representation and navigation in cluttered and dynamic environments.IEEE Robotics and Automation Letters, 6 (3):5081–5088, 2021

  24. [32]

    Zhang, J

    C. Zhang, J. Jin, J. Frey, N. Rudin, M. Mattamala, C. Cadena, and M. Hutter. Resilient legged local navigation: Learning to traverse with compromised perception end-to-end. In2024 IEEE International Conference on Robotics and Automation (ICRA), pages 34–41. IEEE, 2024

  25. [33]

    T. He, C. Zhang, W. Xiao, G. He, C. Liu, and G. Shi. Agile but safe: Learning collision-free high-speed legged locomotion. InRobotics: Science and Systems, 2024

  26. [34]

    P. Roth, J. Frey, C. Cadena, and M. Hutter. Learned perceptive forward dynamics model for safe and platform-aware robotic navigation. InRobotics: Science and Systems, 2025

  27. [35]

    S. Chen, M. Yang, H. Mao, J. Zhang, H. Liu, S. He, D. Zhang, Z. Qiu, and C. Zhang. Sea-nav: Efficient policy learning for safe and agile quadruped navigation in cluttered environments, 2026

  28. [36]

    Truong, D

    J. Truong, D. Yarats, T. Li, F. Meier, S. Chernova, D. Batra, and A. Rai. Learning naviga- tion skills for legged robots with learned robot embeddings. In2021 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), pages 484–491. IEEE, 2021

  29. [37]

    T. Miki, J. Lee, J. Hwangbo, L. Wellhausen, V . Koltun, and M. Hutter. Learning robust percep- tive locomotion for quadrupedal robots in the wild.Science robotics, 7(62):eabk2822, 2022

  30. [38]

    P. Long, W. Liu, and J. Pan. Deep-learned collision avoidance policy for distributed multiagent navigation.IEEE Robotics and Automation Letters, 2(2):656–663, 2017

  31. [39]

    P. Long, T. Fan, X. Liao, W. Liu, H. Zhang, and J. Pan. Towards optimally decentralized multi-robot collision avoidance via deep reinforcement learning. In2018 IEEE international conference on robotics and automation (ICRA), pages 6252–6259. IEEE, 2018

  32. [40]

    Andrychowicz, F

    M. Andrychowicz, F. Wolski, A. Ray, J. Schneider, R. Fong, P. Welinder, B. McGrew, J. To- bin, O. Pieter Abbeel, and W. Zaremba. Hindsight experience replay.Advances in neural information processing systems, 30, 2017

  33. [41]

    W. Ding, S. Li, H. Qian, and Y . Chen. Hierarchical reinforcement learning framework towards multi-agent navigation. In2018 IEEE international conference on robotics and biomimetics (ROBIO), pages 237–242. IEEE, 2018

  34. [42]

    Wijmans, M

    E. Wijmans, M. Savva, I. Essa, S. Lee, A. S. Morcos, and D. Batra. Emergence of maps in the memories of blind navigation agents.AI Matters, 9(2):8–14, 2023

  35. [43]

    F. Xia, A. R. Zamir, Z. He, A. Sax, J. Malik, and S. Savarese. Gibson env: Real-world per- ception for embodied agents. InProceedings of the IEEE conference on computer vision and pattern recognition, pages 9068–9079, 2018

  36. [44]

    A. X. Chang, A. Dai, T. A. Funkhouser, M. Halber, M. Nießner, M. Savva, S. Song, A. Zeng, and Y . Zhang. Matterport3d: Learning from rgb-d data in indoor environments.2017 Interna- tional Conference on 3D Vision (3DV), pages 667–676, 2017

  37. [45]

    E. C. Tolman. Cognitive maps in rats and men.Psychological review, 55(4):189, 1948. 11

  38. [46]

    O’keefe and L

    J. O’keefe and L. Nadel. Pr ´ecis of o’keefe & nadel’s the hippocampus as a cognitive map. Behavioral and Brain Sciences, 2(4):487–494, 1979

  39. [47]

    M. Liu, M. Zhu, and W. Zhang. Goal-conditioned reinforcement learning: Problems and solutions.arXiv preprint arXiv:2201.08299, 2022

  40. [48]

    Schulman, F

    J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov. Proximal policy optimization algorithms.arXiv preprint arXiv:1707.06347, 2017

  41. [49]

    R. Avni, Y . Tzvaigrach, and D. Eilam. Exploration and navigation in the blind mole rat (spalax ehrenbergi): global calibration as a primer of spatial representation.Journal of Experimental Biology, 211(17):2817–2826, 2008

  42. [50]

    R. A. Epstein, E. Z. Patai, J. B. Julian, and H. J. Spiers. The cognitive map in humans: spatial navigation and beyond.Nature neuroscience, 20(11):1504–1513, 2017

  43. [51]

    J. Long, Z. Wang, Q. Li, L. Cao, J. Gao, and J. Pang. Hybrid internal model: Learning agile legged locomotion with simulated robot response. InInternational Conference on Learning Representations, volume 2024, pages 14084–14100, 2024

  44. [52]

    Mysore, G

    S. Mysore, G. Cheng, Y . Zhao, K. Saenko, and M. Wu. Multi-critic actor learning: Teaching rl policies to act with style. InInternational Conference on Learning Representations, 2022

  45. [53]

    L. Wang, K. Yao, Y . Liu, W. Qin, J. Wu, Z. Sun, and Q. Zhu. Puma: Perception-driven unified foothold prior for mobility augmented quadruped parkour.arXiv preprint arXiv:2601.15995, 2026

  46. [54]

    Zargarbashi, J

    F. Zargarbashi, J. Cheng, D. Kang, R. Sumner, and S. Coros. Robotkeyframing: Learning locomotion with high-level objectives via mixture of dense and sparse rewards. InConference on Robot Learning, pages 916–932. PMLR, 2025

  47. [55]

    Schulman, P

    J. Schulman, P. Moritz, S. Levine, M. I. Jordan, and P. Abbeel. High-dimensional continuous control using generalized advantage estimation.CoRR, abs/1506.02438, 2015

  48. [56]

    R. Tarjan. Depth-first search and linear graph algorithms.SIAM Journal on Computing, 1(2): 146–160, 1972. 12 Appendix TABLE OF CONTENTS A Observation Details B Reward Functions C Critic Details D Training Details E Additional Real-World Deployments A Observation Details Propri...

Pith tools

Reviewed June 27, 2026 · model on record in the stance chip above.