Pith. sign in

REVIEW 9 cited by

Learning Vision-Guided Quadrupedal Locomotion End-to-End with Cross-Modal Transformers

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2107.03996 v3 pith:Z7QMYGE6 submitted 2021-07-08 cs.LG cs.CVcs.RO

classification cs.LGcs.CVcs.RO
keywords locomotionmethodobstaclesproprioceptiveterrainchallengingend-to-endenvironments
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

We propose to address quadrupedal locomotion tasks using Reinforcement Learning (RL) with a Transformer-based model that learns to combine proprioceptive information and high-dimensional depth sensor inputs. While learning-based locomotion has made great advances using RL, most methods still rely on domain randomization for training blind agents that generalize to challenging terrains. Our key insight is that proprioceptive states only offer contact measurements for immediate reaction, whereas an agent equipped with visual sensory observations can learn to proactively maneuver environments with obstacles and uneven terrain by anticipating changes in the environment many steps ahead. In this paper, we introduce LocoTransformer, an end-to-end RL method that leverages both proprioceptive states and visual observations for locomotion control. We evaluate our method in challenging simulated environments with different obstacles and uneven terrain. We transfer our learned policy from simulation to a real robot by running it indoors and in the wild with unseen obstacles and terrain. Our method not only significantly improves over baselines, but also achieves far better generalization performance, especially when transferred to the real robot. Our project page with videos is at https://rchalyang.github.io/LocoTransformer/ .

Discussion (0). Sign in to comment.

Forward citations

Cited by 9 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score.

  1. StairMaster: Learning to Conquer Risky Hollow Stairs for Agile Quadrupedal Robots

    cs.RO 2026-06 conditional novelty 6.0 of 10

    A three-stage RL framework with cross-attention, spatial memory, and depth-noise augmentation lets a Unitree Go2 climb 55° hollow stairs zero-shot from simulation.

  2. StairMaster: Learning to Conquer Risky Hollow Stairs for Agile Quadrupedal Robots

    cs.RO 2026-06 unverdicted novelty 6.0 of 10

    StairMaster trains an RL policy that lets a Unitree Go2 quadruped climb hollow stairs up to 55 degrees via zero-shot sim-to-real transfer using cross-attention, SRU memory, and active-perception rewards.

  3. QuadKAN: KAN-Enhanced Quadruped Motion Control via End-to-End Reinforcement Learning

    cs.RO 2025-08 conditional novelty 6.0 of 10

    A KAN-based spline policy for vision-guided quadruped locomotion improves return, distance, and collision avoidance over MLP baselines in PyBullet simulation.

  4. Generalized Locomotion in Out-of-distribution Conditions with Robust Transformer

    cs.RO 2025-07 conditional novelty 6.0 of 10

    A transformer with body tokenization and consistent dropout generalizes to unseen leg damages and sensor noise while trained on limited dynamics and clean observations.

  5. LadderMan: Learning Humanoid Perceptive Ladder Climbing

    cs.RO 2026-06 unverdicted novelty 5.0 of 10

    A hybrid motion-tracking and imitation-reinforcement pipeline produces a depth-based visuomotor policy that lets humanoids climb varied ladders zero-shot on hardware and perform teleoperated manipulation while climbing.

  6. Global-Local Attention Decomposition for Terrain Encoding in Humanoid Perceptive Locomotion

    cs.RO 2026-05 unverdicted novelty 5.0 of 10

    GLAD decomposes terrain encoding via coarse-to-fine attention on elevation maps to separate broad awareness from precise foothold selection in perceptive humanoid locomotion.

  7. Now You See That: Learning End-to-End Humanoid Locomotion from Raw Pixels

    cs.RO 2026-02 unverdicted novelty 5.0 of 10

    An end-to-end policy learns robust humanoid locomotion directly from noisy depth images via high-fidelity sensor simulation, vision-aware distillation from privileged maps, and terrain-specific multi-critic reward shaping.

  8. DPL: Depth-only Perceptive Humanoid Locomotion via Realistic Depth Synthesis and Cross-Attention Terrain Reconstruction

    cs.RO 2025-10 conditional novelty 5.0 of 10

    Combining a blind-backbone policy, cross-attention terrain reconstruction from depth plus proprioception, and realistic synthetic depth with noise enables depth-only full-sized humanoid locomotion over stairs, slopes,...

  9. Hierarchical Reinforcement Learning and Value Optimization for Challenging Quadruped Locomotion

    cs.RO 2025-06 conditional novelty 4.0 of 10

    A hierarchical quadruped controller uses online optimization over the low-level policy's value function to choose footstep targets, improving normalized reward and reducing collisions over an end-to-end baseline witho...

Pith tools