{"id":"39b7bf91-b9de-4285-805d-6598847c5931","arxiv_id":"2602.18164","paper_version":3,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":7.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"GrandTour releases 49 multi-modal legged-robot missions (>10 km, >5 h) with LiDAR, camera, IMU, depth, proprioception, and mm-level RTK-GNSS/total-station ground truth, plus a 52-method state-estimation benchmark.","lead":"A new public dataset records 49 missions in which a quadruped robot carrying 30+ synchronized sensors walked through Swiss cities, forests, mountains, and industrial ruins. The paper also runs 52 open-source odometry and SLAM methods on six of those missions, giving legged-robot researchers a common benchmark similar to what KITTI provided for self-driving cars.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"GT on GNSS-denied / low-TPS-coverage benchmark missions (ARC-2, ARC-7) is INS-dead-reckoning stitched, validated only against fusion inputs; Table 4's 0.11 m @ 60 s drift exceeds the 1–8 cm ATE separations the 52-method benchmark ranks.","rationale":"The paper's value proposition is a KITTI-class legged dataset: survey-grade GT plus a 52-method benchmark resolving centimeter-scale differences. For that to hold, GT must be accurate well below the method separations on every evaluated segment. The weakest condition is Sec. 4.2's GT on benchmark missions with minimal MS60 coverage and intermittent/no GNSS: ARC-7 and ARC-2. The fusion's own disclosure—drift accumulation during outages, Table 4's 0.11 m @ 60 s bound, Fig. 9's growing uncertainty—shows this is a genuine unquantified region, and the validation against raw TPS is a fit-to-input residual, not an independent error bound. I considered two alternatives: (i) benchmark reproducibility (52 configs, 16 undisclosed excluded methods, no released eval scripts) is real and the reader rightly flags it, but it is secondary: with scripts released, rankings computed on drifting GT would still be unreliable; GT validity is upstream. (ii) the 'largest legged-robot dataset' claim vs SubT-MRS is contestable but turns on a defensible definition (legged-specific sensing, proprioception, legged-only mission structure) and does not affect the dataset's utility. The paper deserves credit where independent support exists: TPS measurements carry 1.5 mm range accuracy where visible, the camera-prism calibration cross-check is concrete (3 mm consistency), Fig. 5's LiDAR-camera overlays are strong calibration evidence, and Sec. 4.2/Table 4 disclose the drift mechanism rather than hiding it. The concern is therefore not a rejection—it is a testable condition, which is exactly what CONDITIONAL means. My recommendation is UNCHANGED: the reader's CONDITIONAL verdict stands, with the GT-drift issue as the primary condition. (Minor note: the reader's 42.0% MS60 figure for ARC-2 differs from Table 3's 27.3%; the argument is stronger with the Table 3 value.)","tokens_in":45176,"tokens_out":12690,"duration_ms":97882,"concrete_test":"Leave-out analysis on no-GNSS missions with substantial TPS coverage (CON-1, 76.0% MS60; HAUS-1, 66.0%; or ARC-6, 62.9%): mask all TPS fixes in 15–60 s windows matching ARC-2/ARC-7 outage patterns, re-run Holistic Fusion (released, Nubert et al. 2025), and compare reconstructed positions to the held-out TPS measurements. If drift stays ≲2 cm for the longest realistic outages, the concern is defused; if it reaches 5–10 cm, the ARC-2/ARC-7 rankings and the cm-level benchmark resolution claim must be re-qualified or the benchmark restricted to TPS-covered segments. Secondary check on ARC-7 itself: report per-segment HF-vs-IE deviations, the MS60 outage-duration distribution, and whether published GT spans IE-only coasting segments (i.e., what 'restrict the ground-truth time span' actually excludes).","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—millimeter-level ground truth supporting a 52-method benchmark whose ATE/RTE differences are centimeters—requires GT accurate well below those differences on every evaluated segment. This is least secure exactly where the benchmark is hardest: ARC-7 (GNSS=No, 35.8% MS60 coverage) and ARC-2 (27.3% MS60 coverage, largely indoor; GNSS intermittently unavailable). During TPS occlusions, the Sec. 4.2 Holistic Fusion factor graph (Eq. 1) retains only Inertial Explorer unary poses and the HG4930 IMU—i.e., INS dead-reckoning; the paper itself states this 'lead[s] to drift accumulation,' and Table 4 bounds IE error at 0.01–0.02 m RMS for 10 s outages and 0.11 m for 60 s. Fig. 10's re-lock procedure makes multi-second-to-minute TPS gaps the norm in these missions, and Fig. 9 shows the estimator's own std dev rising when line of sight is lost. The §4.2 validation (HF ATE mean 0.0028 m vs raw TPS) does not bound this drift: TPS positions are factors in the very graph being validated, so small residuals only show self-consistency, not accuracy between fixes. The stated restriction of GT to 'segments with Inertial Explorer pose availability or MS60 measurements' is not a quality gate, since IE emits poses continuously, including GNSS-denied coasting. Consequence: Table 9 shows Coco-LIC and FAST-LIVO2 separated by ~1–8 cm ATE on ARC-2/ARC-7; a 5–10 cm GT drift over 30–60 s outages is at the scale of the rankings, so per-mission ranks (Eq. 2–3) are not trustworthy there. TPS-covered segments and the multi-modal calibration (Fig. 5) are not affected; the concern is specifically the stitched GT on low-coverage missions.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper introduces GrandTour, a 49-mission legged-robotics dataset collected with an ANYmal-D quadruped carrying the Boxi sensor payload. The dataset provides synchronized multi-modal data (three LiDARs, ten cameras, seven IMUs, depth cameras, proprioception, GNSS/INS) and ground-truth trajectories derived from dual-antenna RTK-GNSS with Inertial Explorer post-processing, a Leica MS60 total station, and Holistic Fusion factor-graph integration. The authors also present a localization benchmark on six missions in which 52 open-source odometry/SLAM pipelines are evaluated with ATE/RTE metrics and per-mission ranks. The central claims are that GrandTour is the largest legged-robot dataset to date, that it achieves millimeter-level ground-truth accuracy with sub-millisecond synchronization, and that the accompanying benchmark provides a rigorous cross-method comparison.","tokens_in":45552,"tokens_out":7408,"duration_ms":68327,"significance":"The dataset is potentially a major community resource: it combines a rich sensor suite, substantial environmental diversity, a detailed calibration chain with concrete validation artifacts (3 mm cross-camera prism consistency, LiDAR-camera overlays in Fig. 5), and an unusually detailed evaluation protocol. The open release in both Zarr and ROS formats, with derived outputs and conversion tools, is a genuine strength. If the ground-truth accuracy on low-coverage segments can be rigorously established, this would be a KITTI-class benchmark for quadruped state estimation and perception. However, the current validation of the ground-truth generation is self-referential on exactly the segments where benchmark rankings are decided, which tempers the 'millimeter-level' and '52-method benchmark' claims until the issue is addressed.","major_comments":[{"comment":"The ground-truth validation is self-referential on the segments that matter most. TPS positions are factors in the Holistic Fusion graph (Eq. 1), so the reported 0.0028 m mean ATE against raw TPS measures self-consistency, not absolute accuracy between fixes. When the MS60 line of sight is lost, the graph retains Inertial Explorer unary poses and the HG4930 IMU, i.e., dead-reckoning; the text states that this 'lead[s] to drift accumulation,' and Table 4 gives 0.11 m horizontal position RMS for a 60 s GNSS outage. On benchmark missions ARC-2 (27.3% MS60 coverage), ARC-7 (35.8%, no GNSS), and CON-4 (67.1%, no GNSS), TPS gaps of tens of seconds are the norm due to the re-lock procedure (Fig. 10), so GT on those arcs can plausibly drift to the decimeter level. Since ATE/RTE separations among top LIO/LIVO methods on ARC-2/ARC-7 are about 1–8 cm (Table 9), the per-mission ranks and any 'millim","section":"Sec. 4.2, Eq. (1), Tables 3–4, 8–11, Fig. 10"},{"comment":"The 'millimeter-level accuracy' claim is overgeneralized. The 2.8 mm validation is shown only on the SPX-2 sequence of Fig. 9; the text does not state whether the aggregate numbers cover all missions or a single representative one. No per-mission validation against an independent reference is provided for the 49 missions, many of which have far lower MS60 coverage (e.g., 18.8%, 20.3%, 23.6%, 27.3%). Please replace the global claim with a quantified statement of GT accuracy as a function of coverage, or restrict it to segments where it is actually supported.","section":"Abstract, Sec. 2.3, Sec. 9"},{"comment":"The ordinal benchmark claims would benefit from uncertainty quantification. The text itself notes that many methods are separated by 'a few millimeters to about a centimeter in RTE and a few centimeters in ATE, often comparable to the reported standard deviations,' yet summary ranks (Eq. 3) and statements such as 'Coco-LIC and FAST-LIVO2 obtain the best average ranks' are reported without confidence intervals or significance tests. Given the GT uncertainty on low-coverage missions identified above, rank instability is a real risk. Please provide error bars, bootstrap/permutation intervals, or a sensitivity analysis with respect to GT perturbations.","section":"Sec. 7.1.2, Tables 8–11, Eqs. (2)–(3)"}],"minor_comments":[{"comment":"The list reads '1) State Estimation and localization (Sec. 7.1), 1) Perception (Sec. 7.2), and 3) Locomotion & Navigation'; the second '1)' should be '2)'.","section":"Sec. 7.1 intro"},{"comment":"Typo: 'Intertial Explorer' in Table 2; also 'HF4930 IMU' in Sec. 7.1.2 should be 'HG4930 IMU'.","section":"Table 2, Sec. 7.1.2"},{"comment":"The GNSS column uses 'Yes/No/Partially' but Fig. 7 only distinguishes Yes/No. Please define 'Partially' and explain how it is counted in the figure.","section":"Table 3 / Fig. 7"},{"comment":"Terminology is inconsistent: 'dual RTK-GPS' (Sec. 2.3), 'dual-antenna RTK GNSS' (Sec. 3.1), and 'NovAtel SPAN CPT7' should be unified to avoid confusion about the GNSS receiver/INS.","section":"Sec. 4.2"}],"recommendation":"major_revision","confidential_remarks":"The main barrier is the ground-truth validation gap on low-coverage missions; this is fixable with re-analysis, explicit GT uncertainty traces, and softened claims. I am not questioning the authors' intent, and the dataset/calibration work is substantial. If the authors gate the benchmark to intervals with bounded TPS/GNSS staleness or provide independent validation of the drift behavior, I would be comfortable moving toward acceptance."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is the dataset paper the legged-SLAM community has been waiting for, and the authors earned the milestone. The GT validation, however, has a self-consistency problem that makes the centimeter-level benchmark ranking on two of the six missions less trustworthy than claimed.\n\nWhat's new: 49 missions on an ANYmal-D with a genuinely rich payload — 3 LiDARs, 10 cameras, 8 IMUs, depth, full proprioception — plus RTK-GNSS and Leica total-station reference. No prior legged set has this combination, and the environment spread (alpine, urban, industrial, forest, indoor) is right. The calibration section is concrete: 3-mm cross-camera prism consistency, LiDAR-camera overlays validated in motion, documented time sync. The 52-method benchmark, with per-mission ATE/RTE and a transparent ranking protocol, is a real service. The paper is also unusually honest about tuning sensitivity and mission-dependent performance.\n\nWhere it gets soft: the ground truth is generated by Holistic Fusion, which ingests TPS positions as factors. Validating Holistic Fusion against those same TPS positions (2.8 mm ATE) measures self-consistency, not absolute accuracy. That is fine for TPS-covered segments, where raw fixes bound the error. It is not fine for ARC-2 and ARC-7, where TPS coverage is 27–36% and GNSS is absent or partial. Between fixes, the GT is essentially Inertial Explorer plus IMU dead-reckoning, and Table 4 gives 0.11 m position RMS for a 60 s outage. The paper's Fig. 9 shows the estimator's own std dev rising during line-of-sight loss. The stated restriction to 'segments with Inertial Explorer pose availability or MS60 measurements' is not a quality gate, because IE emits continuously even when coasting. On those missions, the top methods are separated by 1–8 cm, so a 5–10 cm drift in the reference can reorder the leaders. This does not invalidate the dataset — the raw TPS fixes and calibration artifacts remain — but the per-mission ranks on ARC-2/ARC-7 should carry an uncertainty caveat, and the authors should provide a segment-level quality flag or release GT covariance.\n\nTwo smaller items: the 'largest legged-robot dataset' claim sits oddly next to SubT-MRS's 300 km, which includes legged platforms; a legged-only comparison would sharpen the claim. And the benchmark is not fully reproducible as shipped: 16 excluded methods are unnamed and evaluation scripts are not released. Both are fixable.\n\nBottom line: this is a serious, well-documented dataset that deserves careful peer review. If the authors address the GT uncertainty on low-coverage missions — even with an honest per-segment error bound — it will become a default benchmark.","headline":"GrandTour is the legged-SLAM dataset the community needs, but the ground-truth validation on GNSS-denied missions is self-consistency, not absolute accuracy, so the benchmark rankings on ARC-2/ARC-7 inherit unquantified drift.","tokens_in":46220,"tokens_out":2225,"would_cite":true,"duration_ms":20030,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"GrandTour is the largest public legged-robot dataset to date, pairing 49 missions with survey-grade ground truth and a 52-method benchmark.","keywords":["legged robotics","multi-modal dataset","state estimation","SLAM benchmark","multi-sensor fusion","LiDAR-inertial odometry","visual-inertial odometry","ground truth"],"falsifier":"Take the missions with no GNSS and less than half total-station coverage, re-derive the reference trajectory using only the raw total-station fixes with no inertial propagation, and recompute ATE/RTE for the top-ranked methods; if the rank order changes materially beyond a few centimeters, the ground-truth accuracy and the resulting rankings on those missions are falsified.","tokens_in":44955,"feed_emoji":"🤖","tokens_out":5961,"duration_ms":55259,"temperature":0.7,"pith_summary":"GrandTour sets out to give legged-robotics research what large earlier datasets gave driving and aerial vehicles: a shared, open collection of real-world, time-synchronized sensor data plus reference trajectories accurate enough to benchmark state-estimation methods. The paper claims it is the largest legged-robot dataset to date—49 missions, more than ten kilometers of walking, and over five hours of data, spanning alpine, forest, urban, industrial, indoor, and underground environments. The dataset pairs multiple LiDARs, ten cameras, depth cameras, eight IMUs, and full joint/proprioceptive feedback with ground truth from a robotic total station and RTK-GNSS, fused to millimeter-level accuracy. On six representative missions the authors evaluate 52 open-source odometry and SLAM pipelines, reporting per-mission ATE/RTE, ranks, and failures. If the claim holds, this becomes a reference benchmark for legged state estimation, multi-modal fusion, and perception in hard outdoor and indoor conditions.","feed_headline":"Largest legged-robot dataset to date ships with 52-method benchmark","feed_subtitle":"Test SLAM, LiDAR-inertial, and visual-inertial algorithms against millimeter-true trajectories from 49 real-world missions.","key_machinery":"The load-bearing mechanism is the sensor payload plus its calibration and time-synchronization chain: all sensors share a common time source (sub-millisecond for most streams), extrinsics are calibrated to 0.05 mm mechanical tolerance and validated by point-cloud-to-image overlays, and ground-truth poses are generated by factor-graph fusion of total-station 20 Hz positions, post-processed GNSS/INS poses, and IMU measurements—producing a 20 Hz reference trajectory that follows the total station when line of sight exists and bridges occlusions with inertial/GNSS propagation.","core_discovery":"The central claim is that GrandTour is the largest, most comprehensively instrumented public legged-robot dataset to date, and that its survey-grade reference trajectories make it a trustworthy benchmark. The supporting discovery is that centimeter-to-millimeter-level ground truth can be produced under real-field legged locomotion by combining total-station prism tracking, post-processed GNSS/INS, and a high-grade IMU in a factor-graph fusion, and that this reference is good enough to separate 52 open-source odometry/SLAM systems into meaningful rankings with clear failure cases.","pith_inferences":["If GrandTour is adopted the way earlier large datasets were, it could become the default comparison point for legged state estimation, shifting the field's focus from single-metric gains toward cross-mission robustness, initialization, and failure recovery.","The millimeter-level ground-truth claim is only as strong as the reference in GNSS-denied, line-of-sight-blocked segments; an independent check there, such as loop-closure-based map consistency or total-station-only interpolation, would show whether the published rankings survive.","Because the suite includes cross-view images and dense geometry from a moving quadruped, it is a natural testbed for learned depth, relocalization, and neural scene representation, going beyond the paper's own state-estimation benchmark.","A direct stress test would be to run the same benchmark on the lowest-coverage missions with a reference re-derived without dead-reckoning; if method rankings change drastically, the dataset's ranking protocol should be refined."],"forward_implications":["Odometry and SLAM researchers get a common legged-robot test bed with 52 pre-run baselines, per-mission ATE/RTE ranks, and documented failure cases, so new methods can be compared without re-tuning every competitor.","Multi-modal fusion can be studied under real legged dynamics—foot contacts, slipping, body orientation changes—with synchronized LiDAR, camera, depth, IMU, and joint-encoder data on a single platform.","The benchmark's finding that no method dominates across missions, while visual-inertial systems fail most often on dark or featureless sequences, argues for evaluation protocols that emphasize robustness and recovery rather than mean error alone.","Intermediate outputs such as motion-compensated point clouds, leg odometry, terrain maps, and occupancy maps lower the entry barrier for perception, locomotion, and navigation research that does not want to build a full SLAM front end.","Survey-grade total-station ground truth, synchronized to sub-millisecond accuracy, makes it possible to test whether claimed improvements of a few centimeters are real or within reference noise."],"fun_headline_variants":["GrandTour dataset: 49 missions, 52 SLAM methods benchmarked","Leaping benchmark: GrandTour puts 52 odometry systems to test","GrandTour: largest legged-robot dataset with 52-method comparison","Multi-modal legged dataset from ANYmal-D: benchmark and launch","Survey-grade ground truth for legged robots: GrandTour dataset"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The reference trajectories are assumed to remain accurate in the portions of a mission where the total station loses line of sight and GNSS is absent, because those gaps are bridged by inertial/GNSS dead-reckoning; if that drift is larger than the few-centimeter gaps between competing methods, the benchmark's rankings on those segments are not reliable.","fun_headline_variants_meta":{"raw":{"variants":["GrandTour dataset: 49 missions, 52 SLAM methods benchmarked","Leaping benchmark: GrandTour puts 52 odometry systems to test","GrandTour: largest legged-robot dataset with 52-method comparison","Multi-modal legged dataset from ANYmal-D: benchmark and launch","Survey-grade ground truth for legged robots: GrandTour dataset"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000463,"raw_usage":{"total_tokens":2174,"prompt_tokens":792,"completion_tokens":1382,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":536,"completion_tokens_details":{"reasoning_tokens":1287}},"tokens_in":536,"tokens_out":1382,"duration_ms":9649,"temperature":1.0,"reasoning_tokens":1287,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T21:59:52.130277+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take the missions with no GNSS and less than half total-station coverage, re-derive the reference trajectory using only the raw total-station fixes with no inertial propagation, and recompute ATE/RTE for the top-ranked methods; if the rank order changes materially beyond a few centimeters, the ground-truth accuracy and the resulting rankings on those missions are falsified.","supporting_citations":[],"review_version":1}