{"id":"616dfcf6-6f2c-4b7a-be8f-7f212074bfcf","arxiv_id":"2510.20480","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":10,"one_line_summary":"An adaptive factor-graph method fuses LIO and VIO from heterogeneous robots with LiDAR-based detections, uses Wasserstein-distance weighting, and improves localization accuracy under sensory degradation.","lead":"Robots carrying different sensors (one with LiDAR, others with cameras) can keep each other localized in GPS-denied areas by fusing their odometry and inter-robot LiDAR detections in a factor graph that adaptively downweights unreliable measurements. The paper demonstrates large error reductions on real UGV-UAV and multi-UAV flights, and analyzes which positioning directions become unobservable when sensors degrade.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Eq. (22a) appears to mis-specify the detection-factor Jacobian: under the paper's right-perturbation convention, ∂e/∂ξ_Xint should be [e]_×, not [e]_× R_Xint. The novel interpolation factor is therefore incorrect as published.","rationale":"The reader's CONDITIONAL verdict is appropriate, and my read does not overturn it, but it adds a specific correctness condition that the Wasserstein-calibration concern does not capture. The Jacobian error, if confirmed, means the paper must be revised (Eq. 22a) before the method can be reproduced; the experimental evidence does not by itself establish the validity of the published algorithm. I would not reject outright because the error is local and easily fixable, and the experiments may still support the approach if the implementation is corrected. Hence CONDITIONAL. I do not dispute the reader's Wasserstein concern; I regard the Jacobian issue as more directly load-bearing because it affects the correctness of the central construction rather than the strength of an empirical heuristic.","tokens_in":25374,"tokens_out":20862,"duration_ms":194719,"concrete_test":"Numerically differentiate Eq. (18) with central differences at a state with R_W_Xint = Rz(π/2), t_Y - t_X = (1,0,0), and compare to Eq. (22a) for φ=(ε,0,0), δt=0. If the numerical derivative is [[e]_×, -I] and not [[e]_× R, -I], the published Jacobian is wrong. Then re-run the Outdoor #1 and Indoor #2 optimizations with the corrected Jacobian and with Eq. (22a); if the ATE differences are non-negligible, the reported improvements depend on an implementation that differs from the text.","verdict_should_be":"CONDITIONAL","load_bearing_attack":"Under the right-perturbation convention stated in Eq. (2), for e = [(T_W_{Xint})^{-1} T_W_{Yint}]_tr = R_Xint^T (t_Yint - t_Xint), a perturbation T_W_Xint Exp([φ;δt]) changes e to first order as e + [e]_× φ - δt. Therefore ∂e/∂ξ_Xint = [[e]_×, -I3]. Eq. (22a) instead writes [[R_Xint^T Δt]_× R_Xint, -I3], i.e., it right-multiplies the orientation block by R_Xint. The extra factor is not a convention artifact: it fails a numerical derivative test whenever R_Xint is not identity (e.g., R_Xint = Rz(π/2), φ about the body x-axis). Eq. (22b) for Y is consistent with the simple derivation, making the X term look like a typo rather than an intentional parameterization. But as published, the Jacobian that iSAM2 would use for the central quaternary factor is not the derivative of the error in Eq. (18); the factor graph no longer minimizes Eq. (1) as specified. The experimental ATE improvements may come from a corrected implementation, but the manuscript itself does not contain a correct specification of the main novel factor. This is a correctness risk to the central claim, independent of the Wasserstein-calibration and parameter-fitting concerns.","agreement_with_reader":"disagree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes a loosely-coupled factor-graph framework for cooperative localization of a LiDAR-inertial robot X and one or more VIO-equipped robots Y,Z, using 3D LiDAR inter-robot detections. Relative LIO/VIO pose factors are adaptively weighted: LIO degeneracy is detected from the minimum eigenvalue of the scan-matching approximate Hessian, and VIO relative-pose uncertainty is set proportional to the 2-Wasserstein distance between consecutive VIO covariance matrices. A novel quaternary factor interpolates robot poses on SE(3) to fuse asynchronous detections. An observability analysis of a simplified two-pose graph identifies unobservable directions under LIO and VIO degradation. Real-world UGV-UAV and multi-UAV experiments report large ATE reductions (e.g., Outdoor #1 X: 74.591 m to 10.373 m; Indoor #2 Y: 4.671 m to 0.101 m), and two targeted experiments qualitatively confirm the predicted unobservable behavior.","tokens_in":25808,"tokens_out":8009,"duration_ms":73008,"significance":"If the method is correctly specified, the contribution is practically significant: it demonstrates a low-bandwidth, heterogeneous, asynchronous multi-robot localization architecture with explicit degradation awareness, and it provides testable observability predictions that are checked on real data. The paper's strengths are the real-world experimental corpus, the direct experimental tests of the theoretically predicted unobservable directions (Sec. IV-D), and the fact that the observability analysis is a concrete rank/nullspace calculation rather than a heuristic claim. However, the manuscript currently contains a load-bearing Jacobian error in the main novel factor, and the Wasserstein weighting is validated only in-sample with modest correlation in one dataset. Both issues must be addressed before the central claims can be accepted.","major_comments":[{"comment":"Under the right-perturbation convention stated in Eq. (2), for e = [(T_W_Xint)^-1 T_W_Yint]_tr = R_Xint^T (t_Yint - t_Xint), a perturbation T_W_Xint <- T_W_Xint Exp([φ; δt]) gives to first order e -> e + [e]_× φ - δt. Therefore ∂e/∂ξ_Xint = [[e]_×, -I3]. Eq. (22a) instead writes [[R_Xint^T Δt]_× R_Xint, -I3]. The extra R_Xint is not a convention artifact; it fails a finite-difference test whenever R_Xint is not the identity. Since this factor is the paper's main novel fusion element, and since the same expression enters the observability Jacobians (e.g., Eq. (30c) and analogous entries), the manuscript as written does not specify a factor graph that minimizes Eq. (1). Please correct Eq. (22a), re-derive or re-check the Sec. III Jacobians, and provide a numerical-derivative verification of the corrected expressions.","section":"Sec. II-D, Eq. (22a)"},{"comment":"The Wasserstein-scaling factor μ is fitted per robot to ground-truth average error (Table II: μY=260, μZ=500), and Sec. IV-C then reports Pearson correlations between the Wasserstein distance and the relative position error on the same data used to fit/select the scaling and bounds. The correlations range from 0.221 (Indoor #1) to 0.807 (Outdoor #1), so the claim that the Wasserstein distance 'highly correlates with the real-world localization error' (Sec. IV-E) is only partially supported and is in-sample. No ablation is reported with a static VIO covariance or with μ chosen on a training split. Please add a held-out or cross-validation analysis, and quantify the ATE benefit of the Wasserstein weighting compared with a fixed-covariance baseline.","section":"Sec. II-C, Table IV"},{"comment":"Each experimental condition appears to be a single run with no error bars, repeated trials, or statistical tests. The ATE reductions are large and qualitatively convincing, but the word 'significant' in the abstract and conclusion is not supported by repeated measurements. Please add multiple runs for the key degradation experiments (or at least state explicitly that the reported numbers are single-run examples), and report mean/standard deviation or per-run values.","section":"Table III"}],"minor_comments":[{"comment":"The text says the yaw error was reduced 'from 1.668 rad to 0.067 m'; the unit for a rotational ATE should be radians, not meters.","section":"Sec. IV-B, Outdoor #4"},{"comment":"The caption says 'all the presented results were obtained with the same set of parameters,' but the paper later states that OpenVINS parameters were not the same across datasets. Please clarify which parameters are fixed and which are dataset-dependent.","section":"Table II"},{"comment":"The statement that 'utilizing a static covariance matrix ... resulted in loss of the estimate' is anecdotal. If this comparison exists, please report the quantitative ATE/association failure for the static-covariance baseline.","section":"Sec. IV-C"},{"comment":"The notation Aint in Eq. (23a) is used without a formal definition of A. Define it explicitly (e.g., A ∈ {X,Y}) to avoid confusion with the skew-symmetric matrix notation.","section":"Sec. II-D, Eq. (23)"},{"comment":"The phrase 'assumes maximal correlation between the two covariance matrices' is imprecise for the Wasserstein distance. The optimal transport coupling is not the same as maximal correlation; consider rephrasing or citing the closed form more carefully.","section":"Sec. II-C, Eq. (11)"}],"recommendation":"major_revision","confidential_remarks":"The Jacobian error in Eq. (22a) is the main obstacle; it is localized and likely fixable, so major revision rather than rejection seems appropriate. If the authors can supply a numerical derivative check and a small code snippet for the quaternary factor, that would resolve the concern quickly. The in-sample calibration of the Wasserstein weight should be addressed with an ablation or held-out evaluation."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a genuinely useful engineering paper on cooperative localization for heterogeneous teams, and the improvements it reports are real and large. But there is a concrete math error in the central interpolation factor, and the validation has enough soft spots that the claims need to be scaled back or re-supported.\n\nWhat is new and good: the interpolation-based quaternary factor for asynchronous detections is a sensible idea, and the Wasserstein-distance weighting for VIO covariances is a clever, practical heuristic. The observability analysis is a legitimate contribution: it is a rank/nullspace calculation, not circular, and the two experiments that confirm the predicted unobservable directions are convincing. The overall system — LIO on one robot, VIO on another, LiDAR detections in between — addresses a real gap in heterogeneous multi-robot deployment.\n\nNow the problems. First, I checked Eq. (22a). Under the paper's stated right-perturbation convention, the derivative of e = [(T_W^{Xint})^{-1} T_W^{Yint}]_tr with respect to a perturbation of T_W^{Xint} should be [e]_x, -I3. The paper writes [e]_x R_Xint, -I3. The extra R is not a convention artifact; it fails a numerical derivative test for non-identity rotations. So the paper, as written, does not specify the Jacobian that would make the factor graph minimize the cost it defines. The authors may have implemented it correctly in code, but the manuscript does not contain the correct specification of the main novel factor.\n\nSecond, the evaluation. Each experiment is a single run with no error bars, and several parameters (mu, lambda_thr, covariance bounds, OpenVINS settings) are fit to ground truth, likely on the same data. The ATE is computed after trajectory alignment, which removes global yaw and position drift, so the headline numbers overstate absolute localization accuracy. There is no quantitative comparison against a static-weight baseline, only a mention that the static covariance failed in one case. The Wasserstein correlation table shows the heuristic is not universally strong — 0.221 in one indoor set, 0.807 in the best outdoor set — which makes the fitted mu a real concern.\n\nWho this is for: researchers working on multi-robot SLAM and GNSS-denied operation, especially teams with mixed LiDAR and camera platforms. They will find the system design and the observability analysis useful once the Jacobian is corrected.\n\nRecommendation: this deserves serious peer review — not a desk reject — but I would ask for major revision: fix Eq. (22a), rerun with held-out parameters and repeated trials, report errors without alignment or with clearly stated alignment, and add a static-weights baseline.","headline":"Useful engineering with real experimental gains, but the Jacobian in Eq. (22a) is wrong and the evaluation is thinner than the abstract suggests.","tokens_in":26304,"tokens_out":9849,"would_cite":false,"duration_ms":73648,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that a team of robots with complementary sensors can stay accurately localized without GNSS by fusing LiDAR and camera odometry with inter-robot detections, reweighting each sensor by its current reliability.","keywords":["cooperative localization","multi-robot","factor graph","LiDAR-inertial odometry","visual-inertial odometry","Wasserstein distance","sensor degradation","GNSS-denied"],"falsifier":"Measure, with motion-capture ground truth, the relative position error of a Kalman-filter VIO alongside the 2-Wasserstein distance between its consecutive output covariances across a scene with fluctuating visual texture. If the two stop tracking each other (the paper's own reported correlations range from 0.221 to 0.807), the adaptive weight will over-trust or under-trust the VIO, and the claimed rescue of a degraded camera robot fails; this is directly testable in a new deployment.","tokens_in":25229,"feed_emoji":"🤖","tokens_out":8159,"duration_ms":72193,"temperature":0.7,"pith_summary":"The paper tries to show that a heterogeneous robot team—one robot with a LiDAR, others with cameras—can sustain accurate localization in GNSS-denied environments even when one robot's sensor is failing, as long as the robots can see each other. It fuses each robot's own odometry (LiDAR-inertial for the detecting robot, visual-inertial for the detected robots) with 3D detections of teammates in a sliding-window factor graph. The adaptive core is to detect when the LiDAR odometry is degenerate by inspecting the scan-matching Hessian, and to weight visual-inertial odometry by the Wasserstein distance between consecutive filter covariance matrices, which the paper finds correlates with real relative-position error. On real UGV-UAV and UAV-only datasets, this reduces absolute trajectory error dramatically in degraded scenes—for example, from 74.6 m to 10.4 m for a LiDAR-equipped UAV in an open field, and from 4.7 m to 0.1 m for a camera UAV under artificially blacked-out images. The paper also provides an observability analysis that identifies exactly which directions remain unobservable under LiDAR or visual-inertial degradation, and it confirms those predictions experimentally.","feed_headline":"Cooperative lidar-vision fusion cuts a drone's drift from 74 m to 10 m","feed_subtitle":"When one robot's sensor fails, a teammate's 3D detection keeps the whole team's poses accurate in a shared frame.","key_machinery":"The paper is carried by three mechanisms working inside a sliding-window factor graph. First, an interpolation-based quaternary detection factor connects two temporally adjacent poses of the detecting robot and two of the detected robot, using constant-velocity interpolation on the SE(3) manifold to fuse a 3D detection made at an arbitrary timestamp. Second, LiDAR odometry reliability is binarized by thresholding the minimum eigenvalue of the approximate scan-matching Hessian; when it falls below a threshold, the LiDAR relative-pose covariance is inflated so the graph stops trusting it. Third, the relative-pose covariance of the Kalman-filter visual-inertial odometry is set proportional to t","core_discovery":"The central claim is that a loosely-coupled, degradation-aware factor-graph fusion of LiDAR-inertial odometry, visual-inertial odometry, and 3D inter-robot detections lets a team of robots with complementary sensors localize jointly in a shared frame more accurately than any single odometry source when conditions degrade. The method's novel components are an interpolation-based quaternary detection factor that handles asynchronous measurements, a Hessian-eigenvalue test that switches the LiDAR robot between trusted and untrusted modes, and a Wasserstein-distance-based weighting of visual-inertial relative poses that rises and falls with the estimated uncertainty of the filter. The theoretica","pith_inferences":["The Wasserstein-distance weighting is a general reliability signal for Kalman-filter odometry: it could improve single-robot multi-sensor fusion just as well, because it sidesteps the unknown cross-covariance problem without modifying the odometry filter.","The observability analysis implies an active-control rule the authors leave implicit: during LiDAR degradation, the team should deliberately change the relative bearing between robots, since the unobservable direction lies perpendicular to the robot-robot line and rotates as that line rotates.","The per-robot scale factor is fitted from ground-truth error, so the method is not parameter-free; estimating that constant online from detection-consistency residuals would be a natural extension and would make the adaptivity self-tuning.","Since the detections provide no orientation, the detected robot's yaw can remain unobservable indefinitely; a detection that resolves two points on the teammate, such as a marker pair, would close this gap and is testable with the same factor graph."],"forward_implications":["Under normal operation, the factor graph roughly preserves the better of the two odometries; the detected camera robot's 3D absolute trajectory error drops from about half a meter to around a tenth of a meter indoors.","Under visual-inertial degradation (blacked-out images), the camera robot's 3D error drops from 4.671 m to 0.101 m because the Wasserstein-based weighting downweights the drifting odometry and lets detections anchored by the LiDAR robot take over.","Under LiDAR degradation in a large open field, the LiDAR robot's 3D error drops from 74.591 m to 10.373 m and its yaw error from 1.219 rad to 0.360 rad using a single detected camera robot.","If the LiDAR robot is degraded and only one teammate is detected, one degree of freedom (yaw plus perpendicular translation) stays unobservable and the estimate can drift; detecting two teammates makes the problem fully observable and removes the drift.","The detected camera robot's yaw orientation cannot be corrected by cooperative detections alone, so correcting it requires an extra source of yaw information or the camera robot detecting other robots itself."],"fun_headline_variants":["Cooperative lidar-vision fusion cuts drone drift from 74 to 10 m","Degradation-aware multi-robot localization beats single-sensor drift","When one robot's sensor fails, teammate detections keep poses accurate","Factor-graph fusion of async lidar and vision aids GNSS-denied robots"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The method's visual-inertial adaptivity depends on the assumption that the 2-Wasserstein distance between consecutive VIO output covariances grows in proportion to the actual relative-position error, with one per-robot scale constant fitted from ground truth; if that proportionality breaks in a new environment, the cooperative corrections that should rescue a degrading camera robot become unreliable.","fun_headline_variants_meta":{"raw":{"variants":["Cooperative lidar-vision fusion cuts drone drift from 74 to 10 m","Degradation-aware multi-robot localization beats single-sensor drift","When one robot's sensor fails, teammate detections keep poses accurate","Factor-graph fusion of async lidar and vision aids GNSS-denied robots"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000223,"raw_usage":{"total_tokens":1329,"prompt_tokens":811,"completion_tokens":518,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":555,"completion_tokens_details":{"reasoning_tokens":437}},"tokens_in":555,"tokens_out":518,"duration_ms":4988,"temperature":1.0,"reasoning_tokens":437,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-04T08:25:24.537812+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Measure, with motion-capture ground truth, the relative position error of a Kalman-filter VIO alongside the 2-Wasserstein distance between its consecutive output covariances across a scene with fluctuating visual texture. If the two stop tracking each other (the paper's own reported correlations range from 0.221 to 0.807), the adaptive weight will over-trust or under-trust the VIO, and the claimed rescue of a degraded camera robot fails; this is directly testable in a new deployment.","supporting_citations":[],"review_version":1}