Pith. sign in

Grandtour: A legged robotics dataset in the wild for multi-modal perception and state estimation

9 Pith papers cite this work. Polarity classification is still indexing.

9 Pith papers citing it
abstract

Accurate state estimation and multi-modal perception are prerequisites for autonomous legged robots in complex, large-scale environments. To date, no large-scale public legged-robot dataset captures the real-world conditions needed to develop and benchmark algorithms for legged-robot state estimation, perception, and navigation. To address this, we introduce the GrandTour dataset, a multi-modal legged-robotics dataset collected across challenging outdoor and indoor environments, featuring an ANYbotics ANYmal-D quadruped equipped with the Boxi multi-modal sensor payload. GrandTour spans a broad range of environments and operational scenarios across distinct test sites, ranging from alpine scenery and forests to demolished buildings and urban areas, and covers a wide variation in scale, complexity, illumination, and weather conditions. The dataset provides time-synchronized sensor data from spinning LiDARs, multiple RGB cameras with complementary characteristics, proprioceptive sensors, and stereo depth cameras. Moreover, it includes high-precision ground-truth trajectories from satellite-based RTK-GNSS and a Leica Geosystems total station. This dataset supports research in SLAM, high-precision state estimation, and multi-modal learning, enabling rigorous evaluation and development of new approaches to sensor fusion in legged robotic systems. With its extensive scope, GrandTour represents the largest open-access legged-robotics dataset to date. The dataset is available at https://grand-tour.leggedrobotics.com on HuggingFace (ROS-independent), and in ROS formats, along with tools and demo resources.

fields

cs.RO 9

years

2026 9

representative citing papers

FARM: Find Anything using Relational Spatial Memory

cs.RO · 2026-06-13 · conditional · novelty 6.0

A real-time relational spatial memory that parses object queries into spatial predicates, scores them against per-object 3D Gaussians, and retrieves object instances with substantially higher top-K recall than prior scene-graph or video-VLM baselines.

Iterated Invariant EKF for Quadruped Robot Odometry

cs.RO · 2026-04-16 · unverdicted · novelty 5.0

An IterIEKF algorithm for quadruped odometry, relying on proprioceptive kinematic constraints, outperforms vanilla IEKF and SO(3) Kalman filters in accuracy and consistency on simulations and real datasets.

Multimodal embodiment-aware navigation transformer

cs.RO · 2026-04-21 · unverdicted · novelty 4.0

ViLiNT improves goal-conditioned navigation success rates by 166% on average over vision-only baselines across simulations and real rover tests by combining multimodal sensing with embodiment-conditioned diffusion trajectories and clearance scoring.

citing papers explorer

Showing 9 of 9 citing papers.