{"id":"6695727a-c3a2-4e37-8cdc-734b0c301905","arxiv_id":"2505.04438","paper_version":2,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":3,"one_line_summary":"A wheel-encoder-plus-gyroscope dead-reckoning system matches or exceeds radar-inertial odometry on a public benchmark at negligible compute, but independent snow data shows radar remains more accurate.","lead":"A wheel encoder and gyroscope, naively integrated, beat every radar-based method on the Boreas driving odometry leaderboard while using far less computing power. But the authors' own new snowstorm data shows the radar method is actually more accurate, so the headline claim goes beyond the evidence.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Boreas leaderboard margin is the only support for 'outperform' and is confounded by encoder-aided ground truth; the paper's own encoder-free dataset (Table III) shows OG loses to DRO-G in every scenario.","rationale":"The reader and I converge on the same weakest link: the Boreas leaderboard is the sole evidence for the outperform claim, and that comparison is contaminated by shared-sensor ground truth. The paper explicitly acknowledges this in Section IV-A.3 and attempts to defuse it by pointing to higher errors on snow-covered sequences, but that correlation does not quantify the bias. The stronger problem is that the paper's own independent dataset, expressly designed to remove the coupling, shows the opposite pattern: DRO-G outperforms OG in all four scenario classes of Table III. That fact is internal to the manuscript, so no external data are needed to see that the phrase 'in most scenarios' is not supported. This does not undermine the paper's real contribution, a cheap, well-calibrated dead-reckoning baseline, a new slip dataset, and a sensible argument about research priorities, so a conditional acceptance requiring reframed claims and released data remains appropriate. I would not reject the paper, but the abstract and Section V should be revised so that 'outperform' refers to nominal, non-slippery conditions and to the Boreas leaderboard only after the ground-truth coupling is quantified. A direct encoder-free reference on the public 14 sequences is the most decisive way to measure the coupling.","tokens_in":13320,"tokens_out":7443,"duration_ms":69244,"concrete_test":"On the 14 Boreas sequences with public ground truth, recompute the OG relative translation error twice: (a) against the official encoder-assisted RTK-GNSS/INS/encoder reference, and (b) against an encoder-free RTK-GNSS/INS-only reference obtained from the same raw GNSS/INS logs. If the median error increase from (a) to (b) is at least 0.06 percentage points, the size of the leaderboard margin, the 0.20% versus 0.26% result lies within the ground-truth coupling band and cannot support the outperform claim. If an encoder-free reference cannot be reconstructed, a paired per-sequence test on Table III (OG24 versus DRO-G across the 15 independent sequences) would settle whether the claim holds when the ground-truth bias is removed.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central 'outperform' claim rests entirely on the Boreas leaderboard result (Abstract; Table I), where OG achieves 0.20% versus DRO-G's 0.26%. Yet Section IV-A.3 concedes that the Boreas ground-truth trajectory is generated with an RTK-GNSS/INS/encoder solution and the wheel encoder is 'weakly used' in the reference, biasing OG's error favorably by an unquantified amount. This is not a minor caveat: the leaderboard margin is only 0.06 percentage points, so even a small encoder contribution to the reference can manufacture the win. All other data in the paper, specifically the authors' independent dataset whose ground truth is generated without the encoder (Section IV-B.1), show the opposite: in Table III, DRO-G beats OG24 in every condition class (Suburbs no-slip 0.18 vs 0.24; Suburbs slip 0.24 vs 0.50; Campus no-slip 0.17 vs 0.18; Campus drift 0.16 vs 1.00). Thus the headline assertion is not merely unverified; it is contradicted by the paper's own bias-free experiments. The defensible finding, that simple dead reckoning is competitive and cheap, survives, but the abstract's stronger claim should be reframed or supported by a quantified ground-truth-coupling analysis.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Odometer-Gyroscope (OG) odometry, the direct integration of wheel-encoder distance and gyro yaw rate (Section III-A), with a stationary-heuristic gyro bias estimate (Section III-B) and a least-squares calibration of wheel radius and wheel-to-GNSS/INS extrinsics (Section III-C, Eq. 7). The authors benchmark OG against radar-based and lidar-based baselines in two settings: the Boreas public leaderboard (Tables I-II) and a new 15-sequence dataset collected in February 2025, which includes snowstorm episodes with deliberately induced slip (Tables III-IV, Figs. 3-5). Importantly, the independent dataset's ground-truth trajectory is generated without the wheel encoder (Section IV-B.1), removing the sensor-to-reference coupling that exists in Boreas. The paper's central claim is that OG odometry outperforms current state-of-the-art radar-inertial SE(2) odometry in most scenarios at a fraction of the computational cost (Table V), with the Boreas result (0.20%% versus 0.26%% translation error) as the headline example. It concludes that the automotive community should redirect effort from nominal-condition odometry to slip handling and localization.","tokens_in":13538,"tokens_out":9315,"duration_ms":88140,"significance":"The paper asks a well-posed and timely question, and it has genuine strengths. The OG formulation is minimal and transparent: all free parameters (wheel radius, extrinsics, gyro bias) are explicitly identified, and the integration equations are simple enough to be reproduced and checked by hand. The independent dataset in Section IV-B.1 is exactly the right experimental control for the question, since it removes the wheel encoder from the ground-truth generation. The slippage data (Table IV, Figs. 3-5) are granular and falsifiable, the computation-time comparison (Table V) is quantified, and the OG21-versus-OG24 comparison makes a practical point about calibration longevity. The stress-test concern lands on a careful reading: the headline 'outperform' claim rests entirely on the Boreas leaderboard margin, which is confounded by an unquantified encoder-aided reference trajectory, and the paper's own bias-free experiments show the opposite ordering. If the abstract were reframed to 'competitive at negligible cost under no-slip assumptions', the contribution would be a valuable negative/positive result for the odometry community; as written, the headline overstates the evidence.","major_comments":[{"comment":"The headline claim that 'OG odometry can outperform current state-of-the-art radar-inertial SE(2) odometry ... in most scenarios' is supported only by the Boreas leaderboard margin (0.20%% versus 0.26%% translation error). Section IV-A.3 concedes that the Boreas reference trajectory was generated using an RTK-GNSS/INS/encoder solution and that the wheel encoder is 'weakly used' in the reference, 'slightly biasing the results in a positive manner', but the bias is never quantified. The margin is 0.06 percentage points, which is the same order of magnitude as the admitted bias, and the calibration in Eq. (7) is fitted against that same encoder-influenced reference on the 14 public sequences, so the coupling can enter both the estimated parameters and the evaluation. The authors should either quantify the encoder's contribution to the reference trajectory (for example, by recomputing the reference on the public sequences with an encoder-free GNSS/INS solution and re-running the leaderboard comparison, or by providing a sensitivity analysis of the 0.20%% versus 0.26%% margin) or remove the outperform claim from the abstract.","section":"Abstract; Table I; Section IV-A.3"},{"comment":"The independent dataset, whose ground truth is generated without the wheel encoder (Section IV-B.1), directly contradicts the abstract's claim. In Table III, OG24's relative translation error is worse than DRO-G's in every scenario category: 0.24%% versus 0.18%% for Suburbs no-slip, 0.50%% versus 0.24%% for Suburbs slip, 0.18%% versus 0.17%% for Campus no-slip, and 1.00%% versus 0.16%% for Campus drift. OG24 beats only Radar-DG among the baselines, and only in the no-slip rows. The prose in Section IV-B.3 says that OG is 'on par' with state-of-the-art methods, which is more accurate, but the abstract and Table I narrative is never reconciled with this table. The paper should state plainly that on the unbiased benchmark, the state-of-the-art radar baseline is more accurate than OG in every category, and the outperform claim should be restricted to the Boreas leaderboard or dropped entirely.","section":"Table III; Section IV-B.3"},{"comment":"The comparative claim 'on par' is asserted without statistical support. Table III reports only means over 2-5 sequences with no standard deviations or per-sequence values for the no-slip categories, so a 0.01 percentage-point gap (Campus no-slip: 0.17%% versus 0.18%%) is indistinguishable from noise, while the 0.06 percentage-point gap (Suburbs no-slip: 0.18%% versus 0.24%%) is reported without any dispersion measure. The authors should provide per-sequence breakdowns and a measure of spread for the no-slip categories, or temper the comparative wording to match what the data can actually support.","section":"Table III; Section IV-B.3"}],"minor_comments":[{"comment":"The table caption contains a typo: 'BOERAS' should read 'Boreas'.","section":"Table II"},{"comment":"There is a mismatched parenthesis in the sentence 'The car slippage of the Campus drift sequence shown in Fig. 3) has a strong, local impact'; the opening parenthesis before 'Fig. 3' is missing.","section":"Section IV-B.3"},{"comment":"The first-order low-pass filter used to update the gyroscope bias estimate is described only in words; please provide the update equation or a citation so that the implementation is fully reproducible.","section":"Section III-B"},{"comment":"The sentence 'there is no rational explanation for why we should use much more complex, expensive, and sometimes less reliable solutions' overreaches the measured data, which show that the radar and lidar baselines retain their accuracy under slip (for example, 0.13-0.16%% on Campus drift versus 1.00%% for OG). A more measured statement would strengthen the discussion.","section":"Section V"},{"comment":"The text in Section IV-B.3 refers to 'the first slip and no-slip sequences', while the caption of Fig. 5 refers to 'the best Suburbs no-slip and Suburbs slip sequences'; please make the choice consistent and ensure the reader knows which sequences are shown.","section":"Figures 4-5; Section IV-B.3"},{"comment":"The paper calls for new benchmark datasets, and its independent slip dataset is the only bias-free evidence for its comparative claims, yet Section VI defers releasing that dataset to future work. Please state the availability status of the dataset explicitly (including any planned release link) in the current version.","section":"Section VI"}],"recommendation":"major_revision","confidential_remarks":"The radar and lidar baselines compared in the independent-dataset experiments (DRO-G, Radar-DG, Lidar-G) come largely from the same research group as the authors, and the Boreas leaderboard numbers are the group's own published results. This is not improper if the implementations are faithful to the cited works, but it is a reason to require extra care and statistical support in the comparative claims. I also note that this is a six-page workshop-format paper whose technical novelty is intentionally minimal; if the target venue is an archival journal, the editor should weigh whether the empirical contribution (including a dataset that is not yet released) meets that venue's novelty bar. The central scientific message, suitably reframed, is worth publishing somewhere appropriate."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Here's my take on 2505.04438. The meat is not the algorithm—the authors admit the encoder-plus-gyroscope integration is textbook. What's new is the disciplined comparison: putting that trivial baseline on the Boreas leaderboard and collecting a small snowstorm dataset with deliberately induced wheel slip, including per-sequence slip magnitudes. The authors are transparent about the main confound in the leaderboard result: in Section IV-A.3 they note that the Boreas ground truth is generated with an RTK-GNSS/INS/encoder solution and that the encoder is 'weakly used,' biasing the OG result favorably. They don't quantify that bias, and the margin they're defending is only 0.20% vs 0.26%—six hundredths of a percent. That's a thin reed for the abstract's 'outperform SOTA radar in most scenarios.'\n\nWhat the paper does well: the independent dataset in Section IV-B removes the encoder from the ground-truth generation, and there the picture changes. Table III shows DRO-G beating OG in every single condition class—e.g., Suburbs no-slip 0.18 vs 0.24, Campus drift 0.16 vs 1.00. The authors are honest about this, but the abstract and Table I lean on the leaderboard result. The compute comparison is solid, and the finding that old calibration parameters barely degrade performance is practically useful.\n\nSoft spots, in order: (1) missing error bars or confidence intervals on all reported errors; (2) unquantified ground-truth coupling for Boreas; (3) code and dataset not yet released, explicitly deferred to future work; (4) no statistical test comparing methods across sequences. None of these are fatal to the paper's main qualitative message—that a simple wheel encoder is surprisingly competitive and that the community should spend more effort on slip and localization rather than nominal odometry—but they do mean the strong version of the claim is not supported.\n\nThe citation pattern is normal; they engage the relevant radar, lidar, and wheel-odometry literature, including several of their own papers, which is expected here. The writing is clear, and the limitations are flagged where they matter.\n\nWho benefits: researchers working on autonomous driving odometry and benchmarking. The new dataset, once released, could anchor future slip-robustness studies.\n\nRecommendation: yes, this deserves peer review—not because the headline assertion is proven, but because the research-priorities argument is worth airing, and the dataset has real value. I'd ask for a revised abstract that matches the independent-data results, a quantitative sensitivity analysis for the encoder-aided ground truth, and a commitment to release the dataset and code.","headline":"Wheel-encoder odometry is a surprisingly strong baseline, but the paper overstates the leaderboard win and its own independent data undercut the abstract's claim.","tokens_in":14141,"tokens_out":2388,"would_cite":true,"duration_ms":21972,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper argues that a wheel encoder plus a yaw gyroscope, integrated with plain dead-reckoning equations, can outperform state-of-the-art radar-inertial odometry on an autonomous-driving benchmark at negligible computational cost, and…","keywords":["odometry","wheel encoder","yaw gyroscope","dead reckoning","autonomous driving","radar odometry","wheel slip","snow driving"],"falsifier":"Re-run OG odometry on the same benchmark sequences using a reference trajectory produced without any wheel-encoder input, for example GNSS/INS only; if the 0.20% relative translation error no longer beats the 0.26% radar-inertial baseline, the headline outperformance claim fails. Independently, on the paper's snowstorm sequences with encoder-free ground truth, check whether no-slip snowy-road segments keep OG error below about 0.5%; if they exceed that, the claim that nominal conditions are handled is weakened.","tokens_in":13038,"feed_emoji":"🚗","tokens_out":8572,"duration_ms":75102,"temperature":0.7,"pith_summary":"This paper asks whether the autonomous-driving community should keep investing in complex odometry systems. It shows that \"Odometer-Gyroscope (OG) odometry\"—direct integration of wheel-encoder ticks and gyroscope yaw rate—achieves a 0.20% relative translation error on a public multi-season driving benchmark, edging out state-of-the-art radar-inertial methods at 0.26% while using under one millisecond of computation per frame. The authors argue that nominal driving odometry is effectively solved by this simple dead reckoning. To stress-test the approach, they collected snowstorm data with deliberate handbrake-induced slippage and found that only substantial slip (relative translation errors near 1%) defeats the method. The paper's conclusion is that research effort should move from generic odometry to slip detection and localization within prior maps.","feed_headline":"Wheel-and-gyro odometry beats radar at a fraction of the cost","feed_subtitle":"Simple encoder integration tops a driving benchmark at 0.20% error, raising the question of what odometry work remains.","key_machinery":"The load-bearing object is the OG dead-reckoning update. Distance traveled is computed as $d_i = 2\\pi r (c_{i+1} - c_i)/N$ from encoder tick counts, and heading is updated by trapezoidal integration of the gyro: $\\Delta\\theta = (\\omega_i + \\omega_{i+1})(t_{i+1} - t_i)/2$. When the vehicle turns, the wheel's displacement is treated as a circular arc: $\\Delta p = (d_i/\\Delta\\theta)[\\sin(\\Delta\\theta),\\, 1-\\cos(\\Delta\\theta)]^\\top$; when $\\Delta\\theta \\approx 0$, the displacement degenerates to $(d_i, 0)^\\top$.\n\nTwo supporting mechanisms make this work: a gyro-bias estimate obtained by averaging measurements while the vehicle is static and updated with a low-pass filter, and an offline calibration that fits the wheel radius and the encoder-to-GNSS/INS extrinsic parameters by least squares against ground-truth velocities. The identity underlying the argument is that all of the benchmark's sophisticated radar pipelines are, in nominal conditions, recovering the same planar motion that a wheel encoder measures directly.","core_discovery":"The central claim is that the simplest possible dead-reckoning pipeline—counting wheel-encoder ticks for forward distance and integrating a yaw gyroscope for heading—is enough to match or beat the best radar-inertial odometry on nominal road driving. On the public benchmark's leaderboard the OG estimate posts a relative translation error of 0.20% versus 0.26% for the best radar-inertial method, with a per-frame execution time under one millisecond against hundreds of milliseconds. Lidar-inertial odometry can be more accurate, but at roughly three orders of magnitude more computation. The paper's own snowstorm experiments show the method degrades only when wheel slip is deliberately induced: no-slip snow-road sequences stay under 0.50% error, while drift sequences with up to 0.65 m/s RMS lateral slip reach approximately 1% to 2% error. The authors therefore claim that for autonomous driving in nominal conditions, the no-slip assumption is not the limiting factor it is often taken to be.","pith_inferences":["The headline lead may be partly an artifact of the benchmark's own ground truth: the wheel encoder is weakly used in the reference trajectory, and on the paper's independent snowstorm data with encoder-free ground truth the wheel method no longer beats the radar baselines. The outperform claim is therefore safest read as benchmark-specific.","The calibration procedure still leans on a high-accuracy GNSS/INS solution; the paper's suggestion that consumer GPS could replace it is plausible but untested and would require continuous radius and bias estimation to hold over months of driving.","A direct extension would be to run the same encoder-gyro integration on any other dataset with wheel encoders and independent ground truth, to see whether the 0.20% result transfers across vehicles, roads, and traffic conditions; the paper currently tests one platform and one benchmark."],"forward_implications":["In nominal road conditions, a wheel encoder and a yaw gyroscope are enough for competitive odometry, so the marginal value of radar or lidar odometry for ordinary passenger-vehicle autonomy drops sharply.","The practical bottleneck becomes slip: the paper's measurements tie OG error to RMS wheel slip, making slip detection and compensation the natural next problem, with direct applications to stability and driver-assistance systems.","Public benchmarks should include off-nominal trajectories—snow, handbrake turns, sustained drifting—because existing large-scale driving datasets never stress the no-slip assumption.","For autonomous driving, research priority should move from unknown-environment odometry to localization within a prior map and map updating, where odometry serves only as a local prior."],"supporting_citations":[{"why":"Supplies the multi-season driving dataset and leaderboard that anchor the head-to-head comparison.","marker":"[43]"},{"why":"Organizes the competition whose four best-performing pipelines serve as the radar baselines.","marker":"[44]"},{"why":"A representative feature-based radar scan-registration baseline among the top leaderboard methods.","marker":"[23]"},{"why":"A radar-inertial semantic odometry baseline among the top leaderboard methods.","marker":"[45]"},{"why":"A continuous-time radar-inertial odometry baseline included in the leaderboard comparison.","marker":"[27]"},{"why":"Provides the Doppler-velocity radar baseline used in the snowstorm experiments.","marker":"[29]"},{"why":"Introduces the direct radar odometry approach underlying the leading radar baseline.","marker":"[34]"},{"why":"Provides the lidar-inertial continuous-time odometry baseline used as the higher-accuracy comparison.","marker":"[13]"}],"fun_headline_variants":["Simple wheel-gyro odometry beats radar on driving benchmark","Wheel-gyro odometry tops radar at 0.20% error, minimal cost","Basic odometry outperforms radar-inertial on leaderboard","Do we need radar? Simple wheel-gyro wins benchmark","Classic odometry rivals radar for a fraction of cost"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The benchmark's ground-truth trajectory is generated with a solution that also uses the wheel encoder, so the comparison may be biased in the wheel-based method's favor; the paper calls this bias slight but never quantifies it.","fun_headline_variants_meta":{"raw":{"variants":["Simple wheel-gyro odometry beats radar on driving benchmark","Wheel-gyro odometry tops radar at 0.20% error, minimal cost","Basic odometry outperforms radar-inertial on leaderboard","Do we need radar? Simple wheel-gyro wins benchmark","Classic odometry rivals radar for a fraction of cost"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000582,"raw_usage":{"total_tokens":2779,"prompt_tokens":1025,"completion_tokens":1754,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":641,"completion_tokens_details":{"reasoning_tokens":1662}},"tokens_in":641,"tokens_out":1754,"duration_ms":12933,"temperature":1.0,"reasoning_tokens":1662,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T23:28:26.623969+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run OG odometry on the same benchmark sequences using a reference trajectory produced without any wheel-encoder input, for example GNSS/INS only; if the 0.20% relative translation error no longer beats the 0.26% radar-inertial baseline, the headline outperformance claim fails. Independently, on the paper's snowstorm sequences with encoder-free ground truth, check whether no-slip snowy-road segments keep OG error below about 0.5%; if they exceed that, the claim that nominal conditions are handled is weakened.","supporting_citations":[{"cited_title":"Boreas: A multi-season autonomous driving dataset,","cited_arxiv_id":null,"evidence_quote":"Supplies the multi-season driving dataset and leaderboard that anchor the head-to-head comparison."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Organizes the competition whose four best-performing pipelines serve as the radar baselines."},{"cited_title":"Lidar-level localization with radar? the cfear ap- proach to accurate, fast, and robust large-scale radar odometry in diverse environments,","cited_arxiv_id":null,"evidence_quote":"A representative feature-based radar scan-registration baseline among the top leaderboard methods."},{"cited_title":"Cfear++: An efficient radar- inertial semantic odometry,","cited_arxiv_id":null,"evidence_quote":"A radar-inertial semantic odometry baseline among the top leaderboard methods."},{"cited_title":"Continuous-time radar- inertial and lidar-inertial odometry using a gaussian process motion prior,","cited_arxiv_id":null,"evidence_quote":"A continuous-time radar-inertial odometry baseline included in the leaderboard comparison."},{"cited_title":"Are doppler velocity measurements useful for spinning radar odometry?","cited_arxiv_id":null,"evidence_quote":"Provides the Doppler-velocity radar baseline used in the snowstorm experiments."},{"cited_title":"Dro: Doppler-aware direct radar odometry,","cited_arxiv_id":null,"evidence_quote":"Introduces the direct radar odometry approach underlying the leading radar baseline."},{"cited_title":"Are We Ready for Radar to Replace Lidar in All-weather Mapping and Localization?","cited_arxiv_id":null,"evidence_quote":"Provides the lidar-inertial continuous-time odometry baseline used as the higher-accuracy comparison."}],"review_version":1}