ForEnt is a new dataset of time-synchronized RGB-D, LiDAR, proprioceptive and third-person video recordings from 69 labeled forest entrapment events collected across eight sites with a quadruped robot.
Grandtour: A legged robotics dataset in the wild for multi-modal perception and state estimation
9 Pith papers cite this work. Polarity classification is still indexing.
abstract
Accurate state estimation and multi-modal perception are prerequisites for autonomous legged robots in complex, large-scale environments. To date, no large-scale public legged-robot dataset captures the real-world conditions needed to develop and benchmark algorithms for legged-robot state estimation, perception, and navigation. To address this, we introduce the GrandTour dataset, a multi-modal legged-robotics dataset collected across challenging outdoor and indoor environments, featuring an ANYbotics ANYmal-D quadruped equipped with the Boxi multi-modal sensor payload. GrandTour spans a broad range of environments and operational scenarios across distinct test sites, ranging from alpine scenery and forests to demolished buildings and urban areas, and covers a wide variation in scale, complexity, illumination, and weather conditions. The dataset provides time-synchronized sensor data from spinning LiDARs, multiple RGB cameras with complementary characteristics, proprioceptive sensors, and stereo depth cameras. Moreover, it includes high-precision ground-truth trajectories from satellite-based RTK-GNSS and a Leica Geosystems total station. This dataset supports research in SLAM, high-precision state estimation, and multi-modal learning, enabling rigorous evaluation and development of new approaches to sensor fusion in legged robotic systems. With its extensive scope, GrandTour represents the largest open-access legged-robotics dataset to date. The dataset is available at https://grand-tour.leggedrobotics.com on HuggingFace (ROS-independent), and in ROS formats, along with tools and demo resources.
fields
cs.RO 9years
2026 9representative citing papers
A learning-free cross-FOV ensemble plus certainty-coupled GICP keeps place recognition and 6-DoF localization usable under extreme FOV asymmetry and wide-to-narrow cross-sensor pairing where prior methods collapse.
A real-time relational spatial memory that parses object queries into spatial predicates, scores them against per-object 3D Gaussians, and retrieves object instances with substantially higher top-K recall than prior scene-graph or video-VLM baselines.
Lifting non-conservative, actuated, and contact-constrained robot dynamics into an exactly symplectic phase-space map yields state-of-the-art out-of-distribution autoregressive rollout error at low parameter and FLOP cost.
Systematic evaluation on an ANYmal D quadruped shows stereo global-shutter cameras outperform monocular, RGB-D, and rolling-shutter setups, while standard inertial integration can degrade vision-based SLAM under legged dynamics.
An IterIEKF algorithm for quadruped odometry, relying on proprioceptive kinematic constraints, outperforms vanilla IEKF and SO(3) Kalman filters in accuracy and consistency on simulations and real datasets.
Ultra-Fusion presents a unified sliding-window estimator for multi-sensor fusion SLAM supporting WIO/VIO/LIO/LVIO with observability-aware initialization, factor-wise reliability scheduling, and online spatiotemporal calibration.
ViLiNT improves goal-conditioned navigation success rates by 166% on average over vision-only baselines across simulations and real rover tests by combining multimodal sensing with embodiment-conditioned diffusion trajectories and clearance scoring.
Benchmark of MUSE, IEKF, and IS on CYN-1 sequence shows similar RPEs, lower ATE for IEKF and IS, and accuracy-latency trade-offs with open-source evaluation code.
citing papers explorer
-
ForEnt: A Multi-Modal Dataset for Characterizing Quadruped Robot Entrapments in Forest Environments
ForEnt is a new dataset of time-synchronized RGB-D, LiDAR, proprioceptive and third-person video recordings from 69 labeled forest entrapment events collected across eight sites with a quadruped robot.
-
G-PROBE: Cross-FOV Place Recognition and Certainty-Coupled Localization for 3D Point Clouds
A learning-free cross-FOV ensemble plus certainty-coupled GICP keeps place recognition and 6-DoF localization usable under extreme FOV asymmetry and wide-to-narrow cross-sensor pairing where prior methods collapse.
-
FARM: Find Anything using Relational Spatial Memory
A real-time relational spatial memory that parses object queries into spatial predicates, scores them against per-object 3D Gaussians, and retrieves object instances with substantially higher top-K recall than prior scene-graph or video-VLM baselines.
-
CaLiSym: Learning Symplectic Dynamics of Real-World Systems through Structured Canonical Lifts
Lifting non-conservative, actuated, and contact-constrained robot dynamics into an exactly symplectic phase-space map yields state-of-the-art out-of-distribution autoregressive rollout error at low parameter and FLOP cost.
-
Sensor Configuration Matters: A Systematic Evaluation of Multimodal SLAM on Quadruped Robots
Systematic evaluation on an ANYmal D quadruped shows stereo global-shutter cameras outperform monocular, RGB-D, and rolling-shutter setups, while standard inertial integration can degrade vision-based SLAM under legged dynamics.
-
Iterated Invariant EKF for Quadruped Robot Odometry
An IterIEKF algorithm for quadruped odometry, relying on proprioceptive kinematic constraints, outperforms vanilla IEKF and SO(3) Kalman filters in accuracy and consistency on simulations and real datasets.
-
Ultra-Fusion: A Resilient Tightly-Coupled Multi-Sensor Fusion SLAM Framework under Sensor Degradation and Spatiotemporal Perturbation for Intelligent Transportation Systems
Ultra-Fusion presents a unified sliding-window estimator for multi-sensor fusion SLAM supporting WIO/VIO/LIO/LVIO with observability-aware initialization, factor-wise reliability scheduling, and online spatiotemporal calibration.
-
Multimodal embodiment-aware navigation transformer
ViLiNT improves goal-conditioned navigation success rates by 166% on average over vision-only baselines across simulations and real rover tests by combining multimodal sensing with embodiment-conditioned diffusion trajectories and clearance scoring.
-
A Proprioceptive-Only Benchmark for Quadruped State Estimation: ATE, RPE, and Runtime Trade-offs Between Filters and Smoothers
Benchmark of MUSE, IEKF, and IS on CYN-1 sequence shows similar RPEs, lower ATE for IEKF and IS, and accuracy-latency trade-offs with open-source evaluation code.