{"id":"faabba86-0f0b-4d3b-956e-c71e2ce847e3","arxiv_id":"1908.08814","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A multi-spectral visual odometry method that uses photometric bundle adjustment on visible and thermal images, with a fixed-baseline constraint, to recover metric scale without stereo matching.","lead":"The paper presents a visual odometry system that fuses visible and thermal infrared images using direct image alignment, avoiding explicit stereo matching. This matters because it gives robots a way to track motion and recover metric scale in smoke, darkness, or other poor lighting conditions.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Metric-scale claim depends on unquantified visible-LWIR calibration and synchronization; ATE evaluation does not isolate scale error.","rationale":"The reader's weakest assumption identifies exactly the load-bearing condition: the fixed baseline extrinsic T_{b,a} and hardware synchronization are the only sources of metric scale in a system that otherwise starts from random-depth monocular initialization. This is not a peripheral implementation detail. Equations (7) and (8) route every LWIR static-stereo residual through T_{b,a}, so a biased extrinsic directly biases the recovered scale. The paper provides no calibration accuracy metric, no synchronization error bound, and no ablation or sensitivity analysis. The experimental section compares against monocular baselines but does not state whether trajectory alignment preserves scale, so the quantitative ATE table cannot be read as evidence for metric-scale correctness. The additional gating statement about Equation (6) is confusing and should be clarified, but it is secondary to the calibration concern. These issues make the central claim conditional rather than established. I am not claiming the method is wrong; I am claiming the paper does not yet supply the evidence needed to verify the metric-scale claim. Therefore the appropriate verdict remains CONDITIONAL, unchanged from the reader's assessment.","tokens_in":10486,"tokens_out":9372,"duration_ms":102521,"concrete_test":"On sq-01 with motion-capture groundtruth, re-run MSVO with the T_{b,a} translation scaled by factors 0.95, 0.97, 0.99, 1.01, 1.03, and 1.05, and with LWIR timestamps shifted by ±1, ±5, and ±10 ms. Plot the recovered trajectory scale and ATE against the perturbation magnitude. If the scale error is not approximately linear in the baseline perturbation, the claimed scale mechanism is not what the paper describes; if it is linear, the unquantified calibration and synchronization accuracy is the decisive unresolved unknown.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim of recovered metric scale rests on the static-stereo residual in Section IV-B2, Equations (6)-(8). Map points parameterized in the visible camera are projected into LWIR frames through the fixed extrinsic T_{b,a}; any bias in the baseline translation magnitude directly scales the recovered world, and any inter-camera timing offset acts as a motion-dependent virtual baseline error. The paper states the cameras are 'hardware synchronized' (Section V-A) but reports no calibration reprojection error, no extrinsic uncertainty, and no synchronization jitter. The LWIR camera is rolling-shutter, has a small focus range, and suffers from geometric distortions and motion blur (Section V-A), so timing distortion can be significant. Moreover, Table II reports ATE after 'aligned with the groundtruth trajectories' without specifying whether the alignment is SE(3) or Sim(3); if Sim(3) alignment is used, the metric-scale claim is not actually evaluated by that table. There is also an unresolved gating statement at the end of Section IV-D: 'Equation (6) is not used until the metric scale is recovered,' although Equation (6) is the only scale-carrying residual; the paper never states how Equation (6) is treated in the optimization that supposedly recovers scale.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes MSVO, a visual odometry method that uses a rigidly mounted visible-light camera and a long-wave infrared camera without explicit cross-spectral stereo matching. A visible-camera map point is projected into both LWIR frames through the fixed extrinsic T_{b,a}, so that photometric residuals are computed between temporal views of each camera separately. Sliding-window bundle adjustment optimizes keyframe poses, affine brightness parameters, and inverse depths, and the paper argues that the fixed baseline makes the metric scale observable. Experiments on seven self-recorded indoor/outdoor sequences report ATE against motion-capture ground truth, with qualitative demonstrations of metric-scale recovery, fused multi-spectral point clouds, and handling of uncooled-LWIR NUC corruption. The paper claims accurate odometry and semi-dense metric 3D reconstruction without stereo matching.","tokens_in":10696,"tokens_out":8963,"duration_ms":97656,"significance":"If the claims hold, the paper makes a useful contribution to multi-spectral visual odometry: it avoids brittle cross-spectral feature matching, exploits fixed-baseline geometry through direct alignment, provides metric-scale recovery, and addresses the practical NUC-corruption problem of uncooled LWIR cameras. The cost functions and optimization structure are clearly specified, and the reported ATE numbers indicate that the system functions on the recorded sequences. The observability discussion is an attempt to characterize critical motions, which is valuable for a scale-recovering method. However, the evidence for the central metric-scale claim is currently incomplete: the role of the scale-carrying residual in the optimization is described inconsistently, the calibration and synchronization accuracy of the fixed baseline is not quantified, and the ATE evaluation does not isolate scale error. These issues are fixable and do not invalidate the overall idea, but they need to be addressed before the claims can be accepted.","major_comments":[{"comment":"The statement near the end of Section IV-D that \"Equation (6) is not used until the metric scale is recovered\" contradicts the definition of the full cost in Eq. (13), which sums E_{l,k} over active keyframes, with E_{l,k} defined in Eq. (12) as the sum of E^{i,a}_{l,k} and E^{i,b}_{l,k}. Since Eq. (6) is the only scale-carrying residual introduced in the paper, the manuscript never explains how the metric scale is initially recovered if Eq. (6) is excluded from the optimization. Please clarify the optimization schedule: specify which cost function is used before scale convergence, which cost is used in the convergence test, and exactly when Eq. (6) enters the optimization.","section":"Section IV-D and Eq. (13)"},{"comment":"The observability derivation is too terse and contains notation errors that make it impossible to verify. After defining T_b = [R^T, -R^T t; 0^T, 1], the text states t_b = -R t, but the translation vector of that matrix is -R^T t; the definitions of t'_b, t'^*_b, R', s, and t are not stated precisely. The sentence just before Eq. (10), \"When t_b is on the surface formed by t_b, t'_b, and p_w,\" is self-referential and presumably should refer to a different point. As written, the critical-condition equations (9) and (10) cannot be checked, yet they are the only formal support for the claim that the metric scale is observable in the sliding window. Please rewrite the derivation with consistent notation and a complete argument.","section":"Section IV-B3, Eqs. (9)-(10)"},{"comment":"The metric-scale claim rests on the fixed extrinsic T_{b,a} between the visible and LWIR cameras. A bias in the magnitude of the baseline directly scales the recovered world, and any inter-camera timing offset acts as a motion-dependent baseline error. The paper states that the cameras are hardware synchronized but reports no calibration reprojection error, no extrinsic uncertainty, and no synchronization jitter, despite the LWIR camera being rolling-shutter and subject to geometric distortion and motion blur. Please report quantitative calibration and synchronization accuracy, and validate the recovered scale against an independent distance measurement (e.g., known object sizes or an external reference) rather than relying on the fixed baseline alone.","section":"Section V-A and Eqs. (7)-(8)"},{"comment":"The ATE evaluation is described as aligning the estimation results with the groundtruth trajectories, but the alignment model is not specified. Since monocular ORB-SLAM2 and DSO cannot recover metric scale, a Sim(3) alignment is likely used; if so, the ATE values are invariant to a global scale error and do not evaluate the paper's central metric-scale claim. Please state explicitly whether the alignment is SE(3) or Sim(3), and report per-sequence scale error (e.g., the estimated scale factor after SE(3) alignment) for MSVO. Without this, the quantitative evidence for recovered metric scale is incomplete.","section":"Table II and Section V-B"}],"minor_comments":[{"comment":"\"Li Algebra\" should be \"Lie algebra\" in the sentence introducing se(3).","section":"Section III"},{"comment":"The variable used for the state in Eq. (14) is written as f in se(3)^n x R^m, but the update rule in Eq. (15) acts on SE(3) poses; the domain should be SE(3)^n x R^m.","section":"Section IV-D"},{"comment":"\"Duing NUC\" should be \"During NUC.\"","section":"Section IV-E"},{"comment":"The first column header says \"Method\" but the entries are sequence names; the header should be \"Sequence.\"","section":"Table I"},{"comment":"The sentence describing ATE computation is ambiguous: \"the Euclidean distance is computed between the estimation results and the groundtruth\" should specify that the RMSE is over trajectory positions after alignment, not a single Euclidean distance.","section":"Section V-B"},{"comment":"No error bars or repeated trials are reported for the ATE values in Table II; given the short sequences and handheld device, reporting run-to-run variability would strengthen the comparison.","section":"Section V-B"},{"comment":"The manuscript does not mention whether the dataset or code will be released, which limits reproducibility given that the evaluation is on a self-built dataset.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"For the editor: The manuscript is technically interesting and the experiments show that the pipeline runs, but the central metric-scale claim is under-supported as submitted. The most serious issue is the unresolved role of Eq. (6) in scale recovery: the text says it is not used until after scale recovery, yet the full cost includes it. The observability derivation also needs a careful rewrite. I do not see evidence of circularity or misconduct; the issues are technical completeness and experimental evidence. The self-built dataset is understandable given the lack of multi-spectral benchmarks, but if the dataset is not released, the evaluation is hard to reproduce."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThe headline: this paper has a real idea—use direct photometric bundle adjustment across temporally separated views of a fixed RGB+LWIR rig and let the fixed baseline pull the metric scale out, without ever computing cross-spectral correspondences. That is genuinely new in the multi-spectral VO literature and it is worth reading. The system appears to work on its own seven-sequence dataset, with ATE numbers close to DSO-RGB and much better than thermal-only baselines, and the NUC handling is a sensible practical touch. The related work is fair and the component borrowings from DSO and non-overlapping-FOV methods are acknowledged rather than hidden.\n\nThe soft spots are real but not disqualifying. First, the observability argument in Section IV-B3 is too terse. Equations (9) and (10) are presented as critical conditions without a derivation, and the notation is inconsistent enough that I cannot verify the claim that scale is observable in the sliding window. That section should be rewritten with a proper rank/identifiability argument or a citation to one. Second, the metric-scale claim rests on the static-stereo residual, Equations (6)–(8), and therefore on the accuracy of T_{b,a} and on hardware synchronization. The paper says the cameras are hardware synchronized but reports no calibration reprojection error, no extrinsic uncertainty, and no sync jitter. The LWIR camera has rolling shutter, blur, and distortion, which makes sync error a real concern. Third, Table II reports ATE after alignment with ground truth but does not say whether the alignment is SE(3) or Sim(3). If it is Sim(3), the table does not actually evaluate the metric-scale claim. That needs to be stated.\n\nThere is also a genuine internal tension in Section IV: the paper says Equation (6) is not used until the metric scale is recovered, but Equation (6) is the only scale-carrying residual. The reader is never told how scale is recovered if this residual is excluded. This may be a wording slip, but it sits right at the central claim and needs to be fixed.\n\nI would not call this a desk-reject. The contribution is clear, the experiments are encouraging, and the flaws are fixable. The right referee can push the authors to clarify the scale-recovery mechanism, specify the trajectory alignment, and report calibration/sync quality. If they can do that, this is a solid systems paper for the multi-spectral VO community. I would send it to review.","headline":"A credible direct VO extension to RGB+LWIR with a genuinely new fixed-baseline trick, but the metric-scale claim is under-supported by unspecified trajectory alignment and an under-specified scale-recovery step.","tokens_in":11196,"tokens_out":3065,"would_cite":true,"duration_ms":31765,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"A visible–thermal camera pair can do metric visual odometry with no stereo matching.","keywords":["multi-spectral visual odometry","long-wave infrared","direct image alignment","metric scale recovery","bundle adjustment","semi-dense reconstruction","thermal imaging","stereo matching"],"falsifier":"Record a sequence with motion-capture ground truth, then rerun the pipeline with the extrinsic baseline rotated by a known small angle (say 1 degree) and translated by a few millimeters while keeping all other settings unchanged. If the trajectory error and the reconstructed map scale do not degrade as the perturbation grows, the static-stereo term is not actually carrying the metric-scale recovery described in the paper.","tokens_in":10316,"feed_emoji":"🧭","tokens_out":5848,"duration_ms":56027,"temperature":0.7,"pith_summary":"This paper tries to establish that a visible-light camera and a long-wave infrared camera, rigidly fixed with a known baseline, can serve as a metric visual odometry system without ever finding stereo correspondences between the two spectra. Because the two image types share almost no texture, explicit matching is unreliable; the paper instead runs direct image alignment separately on each camera's temporal stream and couples the streams through the fixed baseline geometry. If correct, this allows visible and thermal images to be fused for localization in low-illumination or smoky conditions, and yields a semi-dense 3D reconstruction with both spectral measurements per point. The paper supports the claim with seven indoor and outdoor sequences, comparing against established monocular and stereo baselines.","feed_headline":"Visible-thermal odometry recovers metric scale without stereo matching","feed_subtitle":"Direct alignment on each spectrum plus a fixed baseline yields metric trajectories and semi-dense 3D maps.","key_machinery":"The load-bearing object is the static-stereo direct image alignment term $E^{i,b}_{l,k}$, which reprojects a map point from the visible camera into the LWIR camera through the fixed extrinsic transformation $T_{b,a}$ and compares photoconsistency between two LWIR frames. This term supplies the metric-scale information that temporal monocular alignment alone cannot provide. The paper also derives two critical-motion conditions, Equations (9) and (10), under which the scale becomes unobservable, and argues that the sliding-window optimization generally avoids them.","core_discovery":"The central discovery is that photometric bundle adjustment over multi-view stereo can replace explicit stereo matching in a multi-spectral rig. Each 3D point is initialized with random depth from one camera and projected into subsequent frames of the same camera (temporal multi-view stereo) and, through the fixed extrinsic transform $T_{b,a}$, into the other camera's frames (static stereo). The photometric residual is evaluated only between images from the same camera, so no cross-spectral similarity is ever required. The fixed-baseline constraint makes the otherwise unobservable metric scale observable, provided the motion avoids two critical conditions; over a sliding window with many points and keyframes those conditions generically fail, and the scale converges. The method also discards LWIR frames corrupted by non-uniformity correction and keeps tracking with the visible camera alone, giving odometry continuity during thermal-camera maintenance.","pith_inferences":["An immediate testable extension is to compute, from Equation (10), the minimal two-frame maneuver that guarantees scale observability, and use it as a startup script for practical deployments.","The method's scale accuracy is only as good as the unquantified extrinsic calibration and synchronization between the two cameras; an ablation that perturbs $T_{b,a}$ would quantify how much of the reported accuracy rests on that assumption.","The same architecture should transfer to other sensor pairs with low texture correlation and a known rigid transform, such as visible plus near-infrared or visible plus event camera, where the static-stereo term would play the same scale-pinning role."],"forward_implications":["A multi-spectral rig can be treated as scale-aware without solving cross-modal correspondence, so operation in darkness, smoke, or glare no longer requires thermal-to-visible feature matching.","Thermal non-uniformity correction outages stop being fatal: corrupted thermal frames are dropped and the visible stream alone carries tracking, with scale held by the fixed baseline.","Because scale converges only during sliding-window optimization, the first few frames of a run are not metrically reliable; the scale becomes trustworthy after a motion that excites the baseline direction.","The semi-dense output includes a metric 3D map with per-point visible and thermal intensities, which the paper notes is not produced by earlier multi-spectral methods."],"supporting_citations":[{"why":"Supplies the direct image alignment formulation with affine brightness correction and the Gauss-Newton sliding-window optimization that the proposed cost functions extend.","marker":"[8]"},{"why":"Provides the random-depth initialization and large-scale direct monocular mapping strategy used before metric scale recovery.","marker":"[9]"},{"why":"Defines the prior multi-spectral stereo odometry approach that depends on explicit stereo matching, the limitation the paper removes.","marker":"[21]"},{"why":"Feature-based monocular and stereo visual odometry baseline used for absolute trajectory error comparison in the experiments.","marker":"[23]"},{"why":"Shows how fixed-baseline stereo direct odometry enforces scale; the proposed method adapts that idea to a multi-spectral pair without cross-modal matching.","marker":"[30]"}],"fun_headline_variants":["Thermal-visible odometry without stereo matching","No stereo matching: metric scale from multi-spectral VO","Direct alignment gives metric scale in visible-thermal VO","Visible and thermal images: odometry without stereo matching","Multi-spectral VO recovers scale without explicit stereo"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The fixed baseline transform $T_{b,a}$ between the visible and LWIR cameras must be accurately calibrated and the two cameras must be hardware synchronized; if that transform is wrong, the recovered metric scale is biased even though the photometric optimization may still look consistent.","fun_headline_variants_meta":{"raw":{"variants":["Thermal-visible odometry without stereo matching","No stereo matching: metric scale from multi-spectral VO","Direct alignment gives metric scale in visible-thermal VO","Visible and thermal images: odometry without stereo matching","Multi-spectral VO recovers scale without explicit stereo"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000199,"raw_usage":{"total_tokens":1350,"prompt_tokens":901,"completion_tokens":449,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":517,"completion_tokens_details":{"reasoning_tokens":373}},"tokens_in":517,"tokens_out":449,"duration_ms":4777,"temperature":1.0,"reasoning_tokens":373,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:28:35.424161+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Record a sequence with motion-capture ground truth, then rerun the pipeline with the extrinsic baseline rotated by a known small angle (say 1 degree) and translated by a few millimeters while keeping all other settings unchanged. If the trajectory error and the reconstructed map scale do not degrade as the perturbation grows, the static-stereo term is not actually carrying the metric-scale recovery described in the paper.","supporting_citations":[{"cited_title":"Engel, V","cited_arxiv_id":null,"evidence_quote":"Supplies the direct image alignment formulation with affine brightness correction and the Gauss-Newton sliding-window optimization that the proposed cost functions extend."},{"cited_title":"Engel, T","cited_arxiv_id":null,"evidence_quote":"Provides the random-depth initialization and large-scale direct monocular mapping strategy used before metric scale recovery."},{"cited_title":"Mouats, N","cited_arxiv_id":null,"evidence_quote":"Defines the prior multi-spectral stereo odometry approach that depends on explicit stereo matching, the limitation the paper removes."},{"cited_title":"Mur-Artal and J","cited_arxiv_id":null,"evidence_quote":"Feature-based monocular and stereo visual odometry baseline used for absolute trajectory error comparison in the experiments."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Shows how fixed-baseline stereo direct odometry enforces scale; the proposed method adapts that idea to a multi-spectral pair without cross-modal matching."}],"review_version":1}