{"id":"c5e4b376-742f-430d-b90e-38eea5f2039c","arxiv_id":"2411.16931","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"On a single 140 second garden drive, COIN-LIO beat LIO-SAM in horizontal accuracy and yaw stability, while both handled vertical motion similarly.","lead":"This paper compares two laser-and-inertia based navigation programs, COIN-LIO and LIO-SAM, on an 890 meter loop drive through the Australian Botanic Garden Mount Annan. The authors report that COIN-LIO keeps better horizontal accuracy over long distances while LIO-SAM drifts more in the final stretch.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Internal inconsistency between drift plots (§5.3) and translation-error box plots (§5.4–5.5) leaves the central claim without a consistent quantitative basis.","rationale":"The reader's verdict is CONDITIONAL, focusing on GNSS ground truth under canopy. That is a legitimate concern for any field experiment, but the manuscript contains a more direct problem: its own quantitative evidence is internally inconsistent. The drift figures (§5.3) are not compatible with the translation-error box plots (§5.4–5.5) under any standard evaluation definition. This is not a matter of external validation but of the paper's reported numbers contradicting one another. Consequently, the central claim—that COIN-LIO is superior in the horizontal plane—does not currently have consistent support. I would therefore move the reader's CONDITIONAL to REJECT: the manuscript should not be accepted until the contradiction is resolved, and the proposed test would determine whether the conclusion survives. I set agreement_with_reader to 'partial' because the ground-truth concern is real and should be addressed, but it is secondary to the fact that the reported metrics themselves do not add up.","tokens_in":118,"tokens_out":7510,"duration_ms":134088,"concrete_test":"Request the estimated and ground-truth trajectories from the authors, as offered in §6.5, and reproduce Figures 9–12 using evo_rpe and evo_ape with the same distance bins (95.22, 190.45, 285.68, 380.91, 476.13 m). Specifically: (1) compute relative translation error for delta=476.13 m on both frameworks; (2) compute global position error at segment endpoints; (3) verify whether a 15 m RPE is compatible with the stated ±1 m drift envelope for COIN-LIO. If the recomputed RPE is 15 m, the drift plots must be restated or reinterpreted; if the RPE matches the drift envelope, the box plots are mislabeled or incorrectly computed. This single check settles whether the quantitative basis for the main conclusion is reproducible.","verdict_should_be":"REJECT","load_bearing_attack":"The central claim that COIN-LIO outperforms LIO-SAM in horizontal accuracy rests on two sets of quantitative results that are not obviously compatible. Section 5.3 states that COIN-LIO's x/y drift stayed within ±1 m (z within ±0.8 m) and that LIO-SAM's x/y drift exceeded −4 m near the end of the 890 m loop. Yet Sections 5.4 and 5.5 report, for both frameworks, translation errors of about 5 m at 95.22 m and 'increasing to 15 meters at 476.13 meters' for COIN-LIO, with LIO-SAM reaching 12–15 m at the same distance. If COIN-LIO's global error never exceeds ~1.3 m in any axis, a relative translation error of 15 m over a 476 m segment cannot be reconciled with standard RPE/ATE definitions unless the drift plots and box plots refer to different quantities, alignments, or distance conventions—none of which the paper states. For LIO-SAM, endpoint drifts of −4 m also strain the reported 12–15 m segment errors. The text says COIN-LIO errors are 'generally lower,' but the quoted values are essentially identical. This is more load-bearing than the GNSS ground-truth question: even with perfect ground truth, the reported evidence does not quantitatively support COIN-LIO's superiority. The paired t-test in §5.6 does not resolve this, as it reports only mean relative errors of 3.5% vs 4.5% with no details on the three repeated runs, effect sizes, or distributional assumptions.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper presents a comparative evaluation of two LiDAR-inertial odometry frameworks, COIN-LIO and LIO-SAM, on a newly recorded dataset from the Australian Botanic Garden Mount Annan. The dataset includes 128-beam LiDAR, IMU, and GNSS ground truth over an approximately 890 m looped trajectory with open, paved, and densely vegetated sections. The authors analyze horizontal and vertical trajectory estimates, position drift, relative translation and yaw errors at different traveled distances, and report a paired t-test as evidence that COIN-LIO outperforms LIO-SAM. The central claim is that COIN-LIO maintains superior accuracy in the horizontal plane and over longer trajectories, while both frameworks perform comparably in the vertical plane.","tokens_in":8872,"tokens_out":4285,"duration_ms":36857,"significance":"If the findings are valid, the paper would provide useful comparative evidence on two state-of-the-art LIO systems in an unstructured natural environment, a setting underrepresented in existing benchmarks. The dataset itself (vehicle-mounted, 128-beam LiDAR, garden environment) complements public datasets such as WildPlaces and MulRan. The trajectory visualizations and qualitative observations about LIO-SAM's late-trajectory divergence are informative. However, the quantitative basis for the central claim is undermined by an internal inconsistency between the drift plots and the reported box plots, and the statistical analysis is reported without sufficient detail. The dataset is proprietary and only available on request, which limits reproducibility. The paper's value is as a case study rather than a general benchmark, and it needs substantial revision before its claims can be accepted.","major_comments":[{"comment":"The drift plots and the box plots are numerically irreconcilable. Section 5.3 states that COIN-LIO's x/y drift remained within ±1 m and z drift within ±0.8 m (Fig. 9), and that LIO-SAM's x/y drift exceeded -4 m near the end (Fig. 10). Yet Sections 5.4 and 5.5 report for both frameworks translation errors around 5 m at 95.22 m and increasing to 12–15 m at 476.13 m (Figs. 11b and 12a). If COIN-LIO's global position error is at most about 1.3 m in any axis, a relative translation error of 15 m over a 476 m segment cannot be reconciled with standard definitions of ATE or RPE unless the two sets of plots measure different quantities, use different trajectory alignment procedures, or use different distance conventions. The manuscript does not define which quantity each figure reports. This inconsistency must be resolved before the central claim that COIN-LIO demonstrates superior horizontal accuracy can be considered supported.","section":"§5.3, §5.4, §5.5"},{"comment":"The statistical test reporting is incomplete. The paired t-test is summarized only by mean values (3.5% vs. 4.5% relative translation error; 0.5° vs. 0.7° yaw error) and p < 0.05 thresholds. No information is given about the sample size (presumably the three repeated runs), the pairing structure, the distribution of differences, effect sizes, or whether multiple comparisons across distances were accounted for. The claim that COIN-LIO 'consistently outperformed' LIO-SAM at all distances is not accompanied by distance-wise statistics. These details are necessary to evaluate whether the observed differences are statistically meaningful.","section":"§5.6"},{"comment":"The ground-truth accuracy claim is not validated under canopy. Section 3.1 says the NovAtel PwrPak7D-E1 GNSS receiver ensures 'centimeter-level accuracy,' but the data were collected in a garden with densely vegetated sections where multipath and occlusion can degrade GNSS accuracy to well below that level. If the ground-truth trajectory itself contains horizontal errors, the relative ranking of the frameworks could be an artifact. The authors should provide quantitative validation of the ground truth in the vegetated segments—for example, by checking residuals at the known loop-closing point or comparing with an independent estimator—and should temper the accuracy claim accordingly.","section":"§3.1"},{"comment":"The configuration of both frameworks is described only qualitatively. The COIN-LIO column-shift calibration and the LIO-SAM loop closure, voxel grid, and scan matching parameters are said to be 'adjusted' or 'set to suitable values,' but no concrete parameter values or sensitivity analysis are provided. Because both algorithms are sensitive to such settings, the absence of this information makes it unclear whether the observed performance difference is intrinsic to the methods or an artifact of the chosen configuration. Please include the exact parameter values and, if possible, an ablation or sensitivity study.","section":"§3.1, §6.4"}],"minor_comments":[{"comment":"Figure 8 subfigures are captioned as 'x-y Plane' and the parenthetical says 'widened vertically for more zoom,' but the plots and text refer to the x-z plane for vertical trajectory analysis. The captions should be corrected to 'x-z Plane.'","section":"Figure 8 captions and §5.2"},{"comment":"The RPG Trajectory Evaluation Toolbox is cited as [8] in §2.2, but §4.4 refers to 'RPG-Trajectory [11]' where reference 11 is the KITTI benchmark paper. The citation numbering is inconsistent and should be corrected throughout.","section":"§2.2 and §4.4"},{"comment":"SegMatch is cited as [9], but reference 9 is the EVO package; the citation for SegMatch should be replaced with the appropriate place-recognition reference.","section":"§2.3"},{"comment":"The sentence beginning 'The RPG Trajectory Evaluation Toolbox... offers a comprehensive framework for trajectory analysis.' is immediately followed by 'Provide detailed metrics...' with an incorrect verb form; change to 'It provides detailed metrics...'.","section":"§2.2"},{"comment":"There are typographical artifacts such as 'F AST-LIO2' and 'T rajectory' with irregular spacing; these should be cleaned up in the final version.","section":"Throughout"}],"recommendation":"major_revision","confidential_remarks":"The internal numerical inconsistency between the drift plots and the translation-error box plots is the most serious issue and should be resolved before the manuscript is reconsidered. The proprietary, on-request dataset also limits reproducibility; the editor may wish to weigh whether this meets the journal's data-availability standards. The paper reads like a preliminary technical report, but the topic is relevant and the case study could be valuable with a corrected and more transparent quantitative analysis."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a case study the field can use — a first head-to-head of COIN-LIO and LIO-SAM on a vegetation-rich garden loop — but the numbers as reported do not hold together. Section 5.3 says COIN-LIO's x/y drift stays within ±1 m and LIO-SAM's x/y drift reaches about −4 m near the end. Sections 5.4 and 5.5 then report translation errors of around 5 m at 95.22 m and 12–15 m at 476.13 m, for both frameworks. Those two sets of numbers cannot both be right. If every point on the estimated trajectory sits within roughly a meter of ground truth, a segment error of 15 m is geometrically impossible; if segment errors genuinely reach 15 m, the drift plots cannot be what they appear to be. The paper never defines which quantity is which, or what alignment was used, and the box plots for the two frameworks look nearly identical, which makes the \"generally lower\" claim and the t-test means (3.5% vs 4.5%) hard to credit as reported.\n\nCredit where it's due: the environment is genuinely underrepresented in benchmarks, the dataset description is adequate (140 s, 890 m loop, OS1-128, GNSS/IMU reference), and the trajectory plots give believable visual evidence that LIO-SAM diverges in the final quarter. The limitations list is honest, and running each framework three times is the right instinct.\n\nSoft spots, in rough order. First, the inconsistency above is load-bearing: it is the quantitative basis for the main claim. Second, the GNSS ground truth under dense canopy is asserted as centimeter-level with no validation; multipath under trees can bias the reference in the same direction as one of the frameworks. Third, the statistics are too thin — three runs, no effect sizes, no distributional checks, just p values. Fourth, the LIO-SAM parameters are \"adjusted to suitable values\" without being stated, and the dataset is proprietary, so reproduction from the paper alone is impossible. Minor: a citation slip, SegMatch attributed to reference [9], which is the EVO package.\n\nWho this is for: practitioners choosing a lidar-inertial system for natural terrain, and anyone working on SLAM evaluation methodology — if only as a cautionary example of metric reporting. The flaw is addressable: recompute the errors with clear ATE/RPE definitions and stated alignment, and the comparison either survives or collapses to \"LIO-SAM diverges at the end,\" which the figures already show. I would send it to review with the expectation of heavy revision, but I would not cite the quantitative claims as they stand.","headline":"Useful first COIN-LIO vs LIO-SAM comparison in a vegetation-rich garden, but the drift and segment-error numbers contradict each other, so the quantitative claims need a rework before citation.","tokens_in":9415,"tokens_out":5819,"would_cite":false,"duration_ms":50949,"reading_group":"maybe","serious_thinker":"unclear","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A looped drive through the Australian Botanic Garden Mount Annan shows COIN-LIO keeping the horizontal trajectory nearly drift-free while LIO-SAM drifts by metres in the final quarter of the path.","keywords":["LiDAR-inertial odometry","SLAM evaluation","natural environment dataset","COIN-LIO","LIO-SAM","loop closure","trajectory drift","yaw error"],"falsifier":"Survey several fixed points along the vegetated final quarter with a total station or an RTK base station, then compare both estimated trajectories to those surveyed checkpoints. If LIO-SAM's x-y drift tracks the GNSS solution's own multipath error rather than the surveyed path, the paper's central comparison collapses; if it tracks away from the surveyed path, the conclusion stands.","tokens_in":8369,"feed_emoji":"🌿","tokens_out":5500,"duration_ms":46228,"temperature":0.7,"pith_summary":"This paper tries to establish how two state-of-the-art LiDAR-inertial odometry systems, COIN-LIO and LIO-SAM, hold up in a natural, unstructured environment rather than the urban or indoor settings where such systems are usually tested. Using a 140-second, roughly 890-metre looped drive through the Australian Botanic Garden Mount Annan, recorded with a 128-beam LiDAR plus GPS and IMU, the authors compare trajectory estimates against GNSS ground truth in the horizontal and vertical planes. They find that both systems track elevation well, but COIN-LIO keeps the horizontal path close to ground truth while LIO-SAM accumulates noticeable x-y drift after about 700 metres and larger yaw errors over longer distances. A sympathetic reader would care because it supplies one of the few head-to-head evaluations in dense vegetation and points to loop-closure robustness as the deciding factor for long natural-environment runs.","feed_headline":"COIN-LIO beats LIO-SAM on a tree-lined 890 m loop","feed_subtitle":"In a botanic-garden test, both handle elevation, but only one keeps the horizontal path true.","key_machinery":"The comparison is carried by a looped natural-environment dataset and two error metrics. The dataset, a 140-second, 9.7 GB recording on a roughly 890-metre path with open grass, paved paths, dense vegetation, and a moving car, is the test object that supplies the loop closures and geometrically challenging stretches. The metrics are the Absolute Trajectory Error, which aligns the whole estimated path to ground truth and averages positional deviation, and the Relative Error, which aligns short sub-trajectory pairs and measures local translation and yaw drift over increasing distances. On top of those, the frameworks themselves are the machinery: COIN-LIO augments LiDAR intensity into image patches and fuses photometric error with point-to-plane registration in an iterated extended Kalman filter, while LIO-SAM couples LiDAR and IMU through a factor graph with scan-to-map matching and loop-closure detection. The paper uses these tools to separate vertical performance, where the IMU dominates, from horizontal performance, where loop-closure quality and scan registration decide the outcome.","core_discovery":"On its own terms, the paper reports that COIN-LIO outperforms LIO-SAM in maintaining trajectory fidelity in the horizontal plane. In plots, COIN-LIO's estimated loop stays near the ground truth, with small deviations in densely vegetated stretches that are corrected at loop closure, while LIO-SAM aligns well early on and then diverges clearly in x and y during the final quarter of the loop. Position-drift curves put COIN-LIO within about plus or minus one metre in x and y and 0.8 metres in z, versus LIO-SAM exceeding four metres of drift in x and y by the end while remaining within about 1.5 metres in z. Relative translation and yaw errors grow with distance for both frameworks, but COIN-LIO's means are lower (about 3.5 percent translation error and 0.5 degrees yaw versus 4.5 percent and 0.7 degrees for LIO-SAM), with the differences statistically significant at the distances tested. The authors attribute the gap to COIN-LIO's continuous-time, intensity-augmented registration and effective loop-closure correction, and suggest LIO-SAM's factor-graph scan-to-map matching is less robust to feature-sparse or repetitive vegetation at this scale.","pith_inferences":["A natural next test would be to re-run LIO-SAM with a larger loop-closure search radius and finer voxel filtering; if its final-quarter drift shrinks, the reported gap is partly a configuration effect rather than an architectural limit.","Because the GNSS ground truth was collected under tree cover without independent validation, the absolute drift numbers should be treated as comparative rather than metrological until surveyed checkpoints confirm the reference trajectory.","The intensity-image mechanism that helps COIN-LIO in this garden may transfer to other feature-poor or repetitively textured scenes such as orchards, tunnels, or snow-covered fields, since the same photometric cues do not depend on geometric structure.","One could quantify how much of the horizontal advantage comes from loop closure versus the continuous-time front end by disabling loop closure in COIN-LIO and measuring the drift curve again."],"forward_implications":["For long, looped routes in vegetation-dense parks or forests, COIN-LIO-style intensity-augmented registration is the better default if heading and horizontal position matter.","LIO-SAM can still be used in such environments, but its loop-closure parameters and detection range would need retuning before trusting the final segment of a long run.","Vertical accuracy is the less discriminating axis: both frameworks stayed within roughly one to one-and-a-half metres in z, so elevation-only tasks would not reveal the difference.","Computational cost does not decide the choice here, since the two frameworks ran in nearly identical time on the same hardware.","The new dataset gives the community a compact, repeatable 890-metre benchmark for natural-environment LiDAR-inertial odometry, even though the data are not yet public."],"supporting_citations":[{"why":"Defines the COIN-LIO method, whose intensity-augmented registration and iterated filtering the paper evaluates.","marker":"[1]"},{"why":"Defines the LIO-SAM baseline, a factor-graph LiDAR-inertial system whose loop closure the paper tests.","marker":"[2]"},{"why":"Gives the underlying FAST-LIO2 architecture that explains COIN-LIO's continuous-time registration behaviour.","marker":"[7]"},{"why":"Supplies the trajectory evaluation metrics and alignment protocol used to compute the reported errors.","marker":"[8]"},{"why":"Provides the alternative evaluation package used as a cross-check during metric computation.","marker":"[9]"},{"why":"Documents an existing handheld natural-environment dataset that the vehicle-mounted garden recording complements.","marker":"[3]"},{"why":"Documents a large-scale urban-suburban dataset that frames the need for a vegetated vehicle-mounted benchmark.","marker":"[4]"},{"why":"Supplies the iterative closest point scan-matching basis for understanding feature-poor vegetation stretches.","marker":"[12]"},{"why":"Supplies the normal distributions transform scan-matching alternative, relevant to loop-closure robustness.","marker":"[13]"}],"fun_headline_variants":["Garden loop: COIN-LIO holds line, LIO-SAM drifts wide","COIN-LIO beats LIO-SAM on a tree-lined path","In botanic garden, COIN-LIO wins horizontal accuracy","COIN-LIO stays truer in the horizontal plane at Mount Annan"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole comparison rests on assuming the GNSS receiver used as ground truth stays centimetre-accurate even under the garden's dense tree cover; if the reference drifts there, the reported horizontal advantage could be an artifact of which system the reference agrees with.","fun_headline_variants_meta":{"raw":{"variants":["Garden loop: COIN-LIO holds line, LIO-SAM drifts wide","COIN-LIO beats LIO-SAM on a tree-lined path","In botanic garden, COIN-LIO wins horizontal accuracy","COIN-LIO stays truer in the horizontal plane at Mount Annan"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000652,"raw_usage":{"total_tokens":3013,"prompt_tokens":991,"completion_tokens":2022,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":607,"completion_tokens_details":{"reasoning_tokens":1940}},"tokens_in":607,"tokens_out":2022,"duration_ms":15146,"temperature":1.0,"reasoning_tokens":1940,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T12:43:10.767817+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Survey several fixed points along the vegetated final quarter with a total station or an RTK base station, then compare both estimated trajectories to those surveyed checkpoints. If LIO-SAM's x-y drift tracks the GNSS solution's own multipath error rather than the surveyed path, the paper's central comparison collapses; if it tracks away from the surveyed path, the conclusion stands.","supporting_citations":[{"cited_title":"Coin-lio: Complemen- tary intensity-augmented lidar inertial odometry","cited_arxiv_id":null,"evidence_quote":"Defines the COIN-LIO method, whose intensity-augmented registration and iterated filtering the paper evaluates."},{"cited_title":"Lio-sam: Tightly-coupled lidar in- ertial odometry via smoothing and mapping,","cited_arxiv_id":null,"evidence_quote":"Defines the LIO-SAM baseline, a factor-graph LiDAR-inertial system whose loop closure the paper tests."},{"cited_title":"Fast- lio2: Fast direct lidar-inertial odometry,","cited_arxiv_id":null,"evidence_quote":"Gives the underlying FAST-LIO2 architecture that explains COIN-LIO's continuous-time registration behaviour."},{"cited_title":"A tutorial on quan- titative trajectory evaluation for visual(-inertial) odometry,","cited_arxiv_id":null,"evidence_quote":"Supplies the trajectory evaluation metrics and alignment protocol used to compute the reported errors."},{"cited_title":"Grupp, EVO: Python package for the evaluation of odometry and SLAM, 2017","cited_arxiv_id":null,"evidence_quote":"Provides the alternative evaluation package used as a cross-check during metric computation."},{"cited_title":"”Wild-places: A large- scale dataset for lidar place recognition in un- structured natural environments.” 2023 IEEE in- ternational conference on robotics and automation (ICRA)","cited_arxiv_id":null,"evidence_quote":"Documents an existing handheld natural-environment dataset that the vehicle-mounted garden recording complements."},{"cited_title":"A method for registra- tion of 3-d shapes,","cited_arxiv_id":null,"evidence_quote":"Supplies the iterative closest point scan-matching basis for understanding feature-poor vegetation stretches."},{"cited_title":"The normal distributions transform: A new approach to laser scan matching,","cited_arxiv_id":null,"evidence_quote":"Supplies the normal distributions transform scan-matching alternative, relevant to loop-closure robustness."}],"review_version":1}