{"id":"3b9fa3c3-3e1f-4ad3-a8da-d3911800e19c","arxiv_id":"2507.04321","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A new multi-lidar dataset benchmarks the dome-shaped Livox Mid-360 against the Livox Avia and Ouster OS0-128, and finds Mid-360 produces the most consistent odometry accuracy.","lead":"The authors collected lidar data from three sensor types, including the dome-shaped Livox Mid-360, and benchmarked five SLAM and three registration algorithms. The dataset and benchmarks aim to help researchers compare lidar platforms, though the dataset is not yet public.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Headline-level result (Mid-360 consistently most accurate) is not yet robust: reference trajectories and PTP synchronization are unquantified, and indoor APE differences between Mid-360 and Ouster are ~1 cm, so time offsets or reference noise could reverse rankings.","rationale":"The paper's contribution is a dataset plus a comparative finding, and the comparative finding is the part a reader would actually use: a low-cost dome lidar consistently matches or beats a much more expensive spinning sensor. The dataset configuration is genuinely new relative to TIERS, GEODE, and CTE-MLO, and the ICP-only evaluation is a reasonable way to isolate lidar geometry from IMU fusion. The internal numbers are broadly consistent, and I found no sign of fabrication or gross computational error. The most load-bearing weakness is the reference side: Section III.C and Section III.F do not quantify MoCap or GNSS-RTK accuracy, do not validate per-sensor PTP synchronization, and do not state how trajectories were aligned before computing APE. Because indoor APE margins are at the centimeter level, these unquantified effects are the same order as the claimed effect. A secondary presentation issue is that Table III shows Ouster beating Mid-360 on FAST_LIO2 and FASTER_LIO for IndoorOffice1, so the text's claim that Mid-360 is best 'particularly under FAST-LIO2' is not literally supported; I treat that as a wording issue rather than the core problem. The proposed latency sweep is a direct, feasible check that would settle whether the ranking is an artifact. Since the reader already assigned CONDITIONAL, and the concern is testable rather than proven fatal, I would keep that verdict unchanged.","tokens_in":11350,"tokens_out":5909,"duration_ms":69227,"concrete_test":"Publish the raw rosbags and a synchronization harness, then perform a latency sweep: for each sensor and each sequence (IndoorOffice1, IndoorOffice2, OutdoorRoad), recompute APE after shifting the lidar odometry timestamps relative to the MoCap/GNSS-RTK reference by ±5, ±10, ±20, and ±50 ms, using the paper's stated protocol. Also recompute with 1–2 cm and 0.5–1° perturbations of the extrinsic calibration. If the Mid-360-vs-Ouster ranking changes within any of these plausible offsets, the headline claim is not established; if the ordering survives the full sweep, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim—that the Livox Mid-360 consistently outperforms the Avia and Ouster sensors across both ICP and SLAM—rests on APE differences reflecting lidar scanning properties. But the evaluation in Section III.F is anchored to MoCap (indoor) and Xsens GNSS-RTK (outdoor) references with no stated uncertainty, no per-sensor time-sync validation, and no explicit trajectory-alignment protocol. The indoor margins are very small: in Table III, FAST_LIO2 on IndoorOffice1 gives Mid-360 0.0451±0.0150 m vs. Ouster 0.0446±0.0298 m, and several other entries differ by less than 1 cm. A constant 5–10 ms synchronization offset at walking speed shifts the reference by ~1–5 cm, which is the same magnitude as the claimed differences. The sensor timestamps are heterogeneous—Ouster supplies per-point ring offsets, Mid-360 supplies per-point UNIX timestamps, and the Avia entry in Section IV.A lists no per-point timestamp—so a single PTP sync cannot be assumed to align all three equally. In addition, the SLAM benchmarks use each lidar's built-in IMU (different models and rates), so Table III also conflates lidar geometry with IMU quality. Without an uncertainty analysis for the references or a public release of the raw bags and tuning parameters, the 'consistent winner' conclusion is not separable from reference and synchronization artifacts.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a new multi-lidar dataset collected on a Unitree B1 platform carrying a Livox Avia, a Livox Mid-360, and an Ouster OS0-128, with MoCap ground truth indoors and GNSS-RTK ground truth outdoors. It benchmarks five lidar-inertial SLAM algorithms and three ICP variants on two indoor sequences and one outdoor road sequence, reporting mean and standard deviation of Absolute Pose Error (APE). The central scientific claim is that the Livox Mid-360, a low-cost dome-shaped solid-state lidar, consistently achieves the highest accuracy and stability across both SLAM and ICP methods, and that the dataset is the first to integrate dome-shaped, solid-state, and spinning lidars on a single ground platform.","tokens_in":11649,"tokens_out":3426,"duration_ms":37743,"significance":"The dataset concept addresses a real gap: no existing public dataset combines a dome-shaped Livox Mid-360 with a limited-FOV solid-state lidar and a high-end spinning lidar on the same platform, and the comparison of ICP variants in an IMU-free setting is potentially useful for practitioners. If the dataset is released and the reference trajectories and time synchronization are validated, the work could provide a solid reference for heterogeneous lidar benchmarking and for selecting low-cost sensors for odometry. The paper also gives credit to related work (TIERS, GEODE, CTE-MLO) and is careful to distinguish its contribution from prior datasets. However, the headline comparative claim currently rests on very small APE differences computed from single runs, with no repeated trials and no uncertainty analysis for the ground-truth references, so the significance is conditional on additional validation.","major_comments":[{"comment":"The reference trajectories are load-bearing for the central comparison, but their uncertainty is not quantified. Section III.C states that MoCap and Xsens GNSS-RTK provide ground truth, and Section III.F states that APE is computed against these references, yet no accuracy numbers, no per-sensor time-synchronization validation, and no trajectory-alignment protocol are given. This matters because the claimed indoor margins are extremely small: in Table III, FAST_LIO2 on IndoorOffice1 gives Mid-360 0.0451 ± 0.0150 m versus Ouster 0.0446 ± 0.0298 m, a difference of 0.5 mm, and several other entries differ by less than 1 cm. A constant synchronization offset of a few milliseconds at walking speed shifts the reference by centimeters, which is the same magnitude as the reported differences. The paper should provide a quantitative accuracy assessment of both reference systems, an explicit time-sync validation per sensor (the heterogeneous timestamps described in Section IV.A make this non-trivial), and a statistical comparison over repeated runs.","section":"Section III.C and III.F, Table III"},{"comment":"The empirical basis for the 'consistently' claim is too thin. Only two indoor sequences and one outdoor sequence were collected, with no repeated trials; the reported standard deviation is computed over poses along a single run, not over independent runs. Consequently, the claim that Mid-360 consistently outperforms the other sensors is not statistically supported. The authors should add multiple trials per condition (or at least per trajectory) and report distributions across runs, or perform significance tests such as paired comparisons across trajectories. Without this, the rankings in Tables III and IV cannot be distinguished from run-to-run variability, especially where differences are sub-centimeter.","section":"Section IV.A and Tables III–IV"},{"comment":"The SLAM benchmark conflates lidar geometry with IMU quality. Section III.D benchmarks FAST-LIO2, FASTER-LIO, S-FAST-LIO, GLIM, and FAST-LIO-SAM using each sensor's built-in IMU, and Section IV.A states that the Livox sensors publish IMU data at 200 Hz while the Ouster publishes at 100 Hz, with different IMU models listed in Table II (IAM-20680HT, BMI088, ICM40609). Because these are lidar-inertial odometry systems, differences in APE can be caused by IMU quality or rate rather than by lidar scanning pattern. The abstract highlights performance 'particularly without an IMU in odometry,' but that condition is only realized in the ICP evaluation, not in the SLAM benchmark. The authors should either run IMU-free variants of the SLAM pipelines or clearly separate the lidar-geometry effect from the IMU effect in the analysis.","section":"Section III.D and IV.A, Table II"},{"comment":"The dataset is a primary contribution, but it is not yet available: Section IV.A says it 'will appear online soon.' Since the paper's novelty and reproducibility depend on the raw bags, extrinsic calibration, PTP configuration, and exact preprocessing parameters, the comparative results cannot currently be verified. The authors should provide a public release or, at minimum, a detailed data card with download links, sensor calibration files, and the exact tuning parameters used for each SLAM and ICP method, before the benchmark claims can be fully assessed.","section":"Section IV.A (dataset availability)"}],"minor_comments":[{"comment":"There are several typos and formatting issues: 'Universtiy' in the author affiliation, 'limitted FoV' in the conclusion, 'ourdoor' in the conclusion, 'UA V' in the abstract and Figure 1, and the reference to 'Section VI presents the experimental results' while the actual sections are IV and V. These should be corrected.","section":"Throughout"},{"comment":"The row for 'Ouster (2020)' lists 'OSDome' but does not clearly indicate whether the dome-shaped column is checked; also, the note about the CTE-MLO dataset should be integrated into the table caption or main text for readability.","section":"Table I"},{"comment":"The APE definition is incomplete: the paper should specify whether the error is computed over full SE(3) poses or only positions, and state the trajectory alignment method (e.g., Umeyama alignment with or without scale) and the timestamp interpolation used when comparing to the ground truth.","section":"Section III.F"},{"comment":"For the Livox Avia point cloud, the listed fields do not include a per-point timestamp, unlike the Mid-360 and Ouster entries; this asymmetry is relevant to the synchronization discussion and should be documented explicitly.","section":"Section IV.A"},{"comment":"The ICP hyperparameters, such as the 20 cm indoor and 60 m outdoor correspondence ranges, are presented as fixed choices without a sensitivity analysis. Since ICP results are known to depend on these thresholds and the three lidars have very different point densities and FoVs, a supplementary experiment varying the thresholds would strengthen the comparison.","section":"Section III.E and Table IV"},{"comment":"The grouped boxplots lack a clear description of what quantity is plotted (per-pose APE across all times? per-script errors?) and how many samples each box contains; the caption should state this and define the whiskers and outliers.","section":"Figure 5"}],"recommendation":"major_revision","confidential_remarks":"The paper is a reasonable dataset-and-benchmark contribution for a robotics venue, and the authors are transparent about the dataset not yet being online. In my view the central comparative conclusion is defensible in principle but needs stronger empirical support before publication: the reference uncertainty, time synchronization, and lack of repeated trials are exactly the kinds of issues that can reverse sub-centimeter rankings. I would encourage the editor to require the dataset release and a revised analysis with uncertainty quantification, rather than reject outright, because the platform and the combination of sensors are genuinely useful to the community."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe headline: a dataset configuration that genuinely fills a gap, but the paper's central claim that the Mid-360 is consistently the most accurate rests on thin evidence. The dataset — three lidar categories (Livox Avia, Mid-360, Ouster OS0-128) on one ground platform with MoCap and GNSS-RTK ground truth — is new. CTE-MLO includes the Mid-360 but only on a MAV, and TIERS lacks it. The paper is honest about that, and Table I makes the comparison clear. That contribution is real.\n\nThe benchmarking is standard but competent: five SLAM pipelines and three ICP variants, with APE computed via EVO. The numbers are internally consistent, and the authors wisely limit ICP evaluation to short outdoor segments. The practical takeaway — a low-cost dome lidar often matches or beats a much more expensive spinning sensor — is useful if the results hold.\n\nThe soft spots are load-bearing. First, the statistical basis: two indoor sequences and one outdoor road, no repeated trials, no significance tests. Many APE differences between Mid-360 and Ouster are under a centimeter (e.g., FAST_LIO2 IndoorOffice1: 0.0451 vs 0.0446). That is within the likely uncertainty of the references and time synchronization. The paper gives no uncertainty analysis for the MoCap or GNSS-RTK references, no per-sensor validation of PTP alignment, and no explicit trajectory-alignment protocol. Second, the SLAM benchmarks use each lidar's built-in IMU, so the comparison conflates lidar geometry with IMU quality. The ICP benchmarks avoid that but cover only short segments. And the dataset is not yet public, so the numbers are unverifiable.\n\nNone of this kills the paper as a dataset release. But the 'consistent winner' narrative in the conclusion overstates what three sequences can support. I'd send it to peer review, with two conditions: the dataset and code must be released, and the authors should add uncertainty bounds on the reference trajectories and a more cautious interpretation of the small APE differences.\n\nWho this is for: SLAM researchers choosing a low-cost lidar, and anyone building multi-lidar benchmarks. It deserves a serious referee.","headline":"A genuinely new multi-lidar dataset configuration, but the central claim about Mid-360's consistent superiority needs stronger statistical, synchronization, and uncertainty evidence.","tokens_in":12213,"tokens_out":2057,"would_cite":false,"duration_ms":21673,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that a low-cost dome-shaped solid-state lidar, the Livox Mid-360, consistently yields the most accurate and stable lidar odometry and point-cloud registration across SLAM and ICP methods, and it introduces a dataset that…","keywords":["lidar dataset","solid-state lidar","spinning lidar","Livox Mid-360","point cloud registration","ICP","SLAM benchmarking","lidar odometry"],"falsifier":"Run the same SLAM and ICP pipelines on the released sequences with an independent, higher-accuracy ground-truth reference (for example, a survey-grade total station or recalibrated motion capture) and recompute the APE rankings: if the Livox Mid-360 is no longer consistently the lowest-error sensor across most methods, or if the ordering shifts when time synchronization offsets are deliberately perturbed, then the paper's central claim would be overturned.","tokens_in":11165,"feed_emoji":"📡","tokens_out":7546,"duration_ms":71358,"temperature":0.7,"pith_summary":"This paper introduces a multi-lidar dataset, collected on a single moving robot, that is the first to include a dome-shaped solid-state lidar (the Livox Mid-360) together with another solid-state lidar (the Livox Avia) and a high-end spinning lidar (the Ouster OS0-128) under synchronized ground truth. On top of the dataset, the authors benchmark five SLAM algorithms and three ICP point-cloud registration variants without IMU assistance, using indoor office sequences with motion-capture ground truth and an outdoor road sequence with GNSS-RTK ground truth. The central result is that the low-cost, dome-shaped Mid-360 delivers the highest accuracy and the tightest error distribution across most SLAM and ICP methods, while the long-range Avia excels only in one outdoor SLAM configuration and the spinning Ouster is competitive but more variable. The practical point the authors are trying to establish is that sensor design—vertical field of view and scan coverage—matters more than price class for odometry in GNSS-denied environments, and that cheap dome lidars can serve as reliable primary sensors.","feed_headline":"Low-cost dome lidar beats high-end spinning sensors in accuracy tests","feed_subtitle":"New dataset places dome, solid-state, and spinning lidars on one robot; the low-cost Mid-360 is most consistent.","key_machinery":"The central object is the new dataset itself: simultaneous, time-synchronized (PTP) point clouds from three lidar types mounted rigidly on a mobile quadruped robot, with motion-capture ground truth indoors and GNSS-RTK ground truth outdoors. What carries the argument is the controlled single-platform configuration, which removes platform and synchronization differences and lets APE differences be attributed to the sensors and the registration methods. On top of that, the standardized evaluation protocol—same SLAM pipelines, same ICP variants, fixed range constraints per environment, and APE as the common metric—is what makes the Mid-360's consistency a measurable, comparative result rather than a single anecdote.","core_discovery":"The paper's central claim is that the Livox Mid-360, a dome-shaped solid-state lidar with a 360° horizontal field of view and roughly 25° vertical coverage, consistently produced the lowest mean Absolute Pose Error (APE) and the most stable results across both SLAM algorithms (FAST-LIO2, FASTER-LIO, S-FAST-LIO, GLIM, FAST-LIO-SAM) and ICP variants (point-to-point KISS-ICP, point-to-plane Open3D-GICP, hybrid GenZ-ICP) in indoor and outdoor tests. In indoor offices, the Mid-360 beat the Ouster OS0-128 and the Livox Avia on most algorithms; outdoors, it remained the most consistent, whereas the Avia, which has the longest range (450 m), won only under FASTER-LIO, and the Ouster, which has the highest resolution, never dominated. The authors attribute the Mid-360's advantage to its dome-shaped design, which provides a wide vertical field of view and dense scan coverage of the immediate surroundings, and they emphasize that this finding holds in IMU-free odometry, isolating the lidar's geometric contribution.","pith_inferences":["If the consistency result generalizes, integrators of indoor service robots could replace high-end spinning lidars with dome-shaped solid-state units, cutting sensor cost sharply without sacrificing odometry accuracy.","Because the environment set is small (two offices, one road), a natural next experiment is to add corridors, vegetation-heavy areas, and weather variation to test whether the Mid-360's geometric advantage persists across scene types.","A direct corollary of the wide-vertical-FoV explanation is that tilting the spinning Ouster to increase vertical coverage should shrink its APE gap to the Mid-360, which can be checked on the released sequences.","The benchmark strips out IMU assistance to isolate sensor geometry; fusing the same data with IMUs in the SLAM pipelines would reveal how much of the dome lidar's advantage survives tight sensor coupling."],"forward_implications":["A low-cost dome-shaped solid-state lidar such as the Livox Mid-360 can serve as a primary odometry sensor in GNSS-denied indoor environments, matching or outperforming spinning sensors that cost several times more.","In IMU-free odometry, scan geometry—vertical field of view and coverage of nearby surfaces—can matter more for registration accuracy than maximum range or nominal resolution.","Narrow-FoV solid-state lidars like the Avia retain value for long-range outdoor SLAM, but their poor local registration means a multi-lidar system pairing them with a dome sensor would exploit both strengths.","The dataset offers a common platform, ground truth, and evaluation protocol that future lidar odometry and registration work can benchmark against, closing the gap left by datasets that exclude dome-shaped sensors."],"supporting_citations":[{"why":"Previous multi-modal lidar dataset that includes spinning and solid-state sensors but omits dome-shaped lidars; establishes the gap the new dataset fills.","marker":"[17]"},{"why":"Existing CTE-MLO dataset that includes the Livox Mid-360 but only on a MAV platform, so it cannot support direct cross-modality comparisons on the same platform.","marker":"[20]"},{"why":"Heterogeneous lidar dataset for degenerate environments that lacks dome-shaped sensors and IMU-free registration tests; applies to the baseline comparison.","marker":"[19]"},{"why":"The classic spinning-lidar autonomous driving benchmark; represents the dominant dataset format that lacks solid-state and dome sensors.","marker":"[13]"},{"why":"FAST-LIO2, the tightly coupled lidar-inertial odometry algorithm used as one of the five SLAM baselines in the benchmark.","marker":"[24]"},{"why":"KISS-ICP, the point-to-point registration pipeline used for the ICP evaluation.","marker":"[29]"},{"why":"Open3D library, used to implement the point-to-plane GICP variant.","marker":"[30]"},{"why":"GenZ-ICP, the hybrid registration pipeline used for the adaptive point-to-point/point-to-plane evaluation.","marker":"[32]"}],"fun_headline_variants":["Dome lidar Mid-360 tops spinning and solid-state rivals in accuracy","Low-cost dome lidar outperforms high-end spinning in SLAM tests","Mid-360 dome lidar most accurate across SLAM and ICP variants","Dome-shaped Livox Mid-360 leads accuracy in multi-lidar benchmark","Low-cost Mid-360 wins on accuracy across lidar types"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The comparison assumes that the motion-capture and GNSS-RTK reference trajectories, plus the PTP time synchronization and extrinsic calibration, are accurate enough that the reported APE differences reflect lidar characteristics rather than time offsets or calibration errors, and that two office sequences and one outdoor road are representative enough to support general conclusions.","fun_headline_variants_meta":{"raw":{"variants":["Dome lidar Mid-360 tops spinning and solid-state rivals in accuracy","Low-cost dome lidar outperforms high-end spinning in SLAM tests","Mid-360 dome lidar most accurate across SLAM and ICP variants","Dome-shaped Livox Mid-360 leads accuracy in multi-lidar benchmark","Low-cost Mid-360 wins on accuracy across lidar types"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000828,"raw_usage":{"total_tokens":3706,"prompt_tokens":1120,"completion_tokens":2586,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":736,"completion_tokens_details":{"reasoning_tokens":2489}},"tokens_in":736,"tokens_out":2586,"duration_ms":17082,"temperature":1.0,"reasoning_tokens":2489,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T19:49:23.423419+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same SLAM and ICP pipelines on the released sequences with an independent, higher-accuracy ground-truth reference (for example, a survey-grade total station or recalibrated motion capture) and recompute the APE rankings: if the Livox Mid-360 is no longer consistently the lowest-error sensor across most methods, or if the ordering shifts when time synchronization offsets are deliberately perturbed, then the paper's central claim would be overturned.","supporting_citations":[{"cited_title":"Multi-Modal Lidar Dataset for Benchmarking General-Purpose Localization and Mapping Algorithms","cited_arxiv_id":"2203.03454","evidence_quote":"Previous multi-modal lidar dataset that includes spinning and solid-state sensors but omits dome-shaped lidars; establishes the gap the new dataset fills."},{"cited_title":"Cte-mlo: Continuous-time and efficient multi-lidar odometry with localizability- aware point cloud sampling, 2025","cited_arxiv_id":null,"evidence_quote":"Existing CTE-MLO dataset that includes the Livox Mid-360 but only on a MAV platform, so it cannot support direct cross-modality comparisons on the same platform."},{"cited_title":"Heterogeneous lidar dataset for benchmarking robust localization in diverse degenerate scenarios, 2024","cited_arxiv_id":null,"evidence_quote":"Heterogeneous lidar dataset for degenerate environments that lacks dome-shaped sensors and IMU-free registration tests; applies to the baseline comparison."},{"cited_title":"Vision meets robotics: The kitti dataset","cited_arxiv_id":null,"evidence_quote":"The classic spinning-lidar autonomous driving benchmark; represents the dominant dataset format that lacks solid-state and dome sensors."},{"cited_title":"Fast-lio2: Fast direct lidar-inertial odometry, 2021","cited_arxiv_id":null,"evidence_quote":"FAST-LIO2, the tightly coupled lidar-inertial odometry algorithm used as one of the five SLAM baselines in the benchmark."},{"cited_title":"Kiss-icp: In defense of point-to-point icp – simple, accurate, and robust registration if done the right way.IEEE Robotics and Automation Letters , 8(2):1029–1036, February 2023","cited_arxiv_id":null,"evidence_quote":"KISS-ICP, the point-to-point registration pipeline used for the ICP evaluation."},{"cited_title":"Genz-icp: Generalizable and degeneracy-robust lidar odometry using an adaptive weighting.IEEE Robotics and Automation Letters , 10(1):152–159, January 2025","cited_arxiv_id":null,"evidence_quote":"GenZ-ICP, the hybrid registration pipeline used for the adaptive point-to-point/point-to-plane evaluation."}],"review_version":1}