{"id":"124260ed-d570-443b-a6af-8ba634871f72","arxiv_id":"2608.12866","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":11,"one_line_summary":"AMR-Pose combines an active red-blue LED marker module with a probabilistic switching PnP estimator to achieve 6-DoF relative pose tracking between AUVs under partial marker visibility.","lead":"This paper presents AMR-Pose, a system that uses four colored LEDs on one underwater robot and a camera on another to estimate their relative position and orientation in murky water. It combines a probabilistic estimator that switches between full and partial marker views, and tests show it tracks accurately when markers disappear briefly.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Constant-acceleration motion model (Eq. 16) is the load-bearing bridge for partial-visibility updates; experiments only test smooth 30° yaw, so robustness to abrupt acceleration is unvalidated.","rationale":"The paper's internal derivation (Sections III-C to III-F) appears mathematically consistent: the Jacobian expressions in Eqs. 18–23, the left-Jacobian treatment, and the pixel-level Jacobian in Eq. 52 all agree with first-order Lie-group EKF conventions. The reported experiments are substantial and the motion-capture ground truth supports the observed 0.032 m/0.032 rad RMSE on the tested trajectory. However, the central claim extends beyond that trajectory: 'robust relative pose estimation under partial observations.' The only mechanism that bridges partial observations is the motion model in Eq. (15)–(16). A constant-acceleration random walk is a strong assumption for a maneuvering AUV. The test suite deliberately uses a slow, smooth yaw maneuver, so it does not exercise model mismatch. This is not a mathematical error but an unvalidated boundary of the claim. The reader identified the same assumption; we agree. Other concerns (lack of code/data, no error bars, undisclosed parameters) are about verification completeness rather than the soundness of the argument, and the reader's conditional verdict already captures them. A targeted simulation or a more dynamic experiment would directly settle whether the motion model is the load-bearing limitation.","tokens_in":16287,"tokens_out":8831,"duration_ms":86920,"concrete_test":"Run the PSwPnP estimator (same equations and noise parameters as in the paper) on a synthetic trajectory with a step change in acceleration, e.g., 0.5 m/s² for 2 s then 0, and evaluate pose RMSE during epochs with only 2–3 visible LEDs. If translation or rotation RMSE exceeds 0.15 m or 0.2 rad (compared with 0.036 m/0.042 rad in the paper's 3-LED segment), the constant-acceleration prior is the load-bearing limitation.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that PSwPnP maintains accurate, smooth pose estimates when only 1–3 LEDs are visible. The load-bearing mechanism is the Lie-group EKF's constant-acceleration random-walk model (Eqs. 15–16): with fewer than four LEDs, each pixel measurement constrains at most two degrees of freedom, so the remaining DOFs are supplied by the predicted pose. The reported experiments only exercise a scripted yaw maneuver with smooth 30° left/right rotations; no step change in acceleration, rapid translation, or high-bandwidth disturbance is tested. Under abrupt leader acceleration, the predicted pose biases the few pixel residuals, and the visible LEDs cannot fully correct the unconstrained DOFs, leading to drift. Since the pose-level PnP update only fires when all four LEDs are visible, the accuracy during the 3-LED segment (e_mid3 = 0.036 m / 0.042 rad) reflects the smoothness prior at least as much as the measurements. Thus the claimed robustness to partial visibility is conditional on the constant-acceleration assumption, and the paper provides no evidence that this condition holds for general AUV maneuvering.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes AMR-Pose, a hardware-plus-estimation system for 6-DoF relative pose estimation between a follower AUV and a leader AUV that carries an active LED marker array (one red central LED and three blue peripheral LEDs). The estimation layer, PSwPnP, is an SE(3) Lie-group EKF with a constant-twist-acceleration motion model, probabilistic marker association, existence-aware visibility management, and measurement updates that switch between a pose-level PnP update when all four LEDs are visible and a pixel-level tight-coupling update when one to three LEDs are visible. The paper reports water-tank experiments with motion-capture ground truth over 45 trials, including ablations, comparisons with frame-wise EPnP-LM and GMLPnP baselines, annotated 4-3-4 visibility transitions, and a closed-loop leader-follower experiment. The headline results are an overall translation RMSE of 0.032 m and rotation RMSE of 0.032 rad for PSwPnP.","tokens_in":16548,"tokens_out":9871,"duration_ms":98879,"significance":"The proposed framework is a reasonable integration of existing techniques: active LED fiducials, Lie-group EKF prediction, probabilistic data association, and visibility-based measurement selection. The experimental methodology has real strengths: external motion-capture ground truth, a dedicated 4-3-4 visibility transition study, and ablation variants that isolate components. The estimator derivation follows standard Lie-group EKF literature and appears internally consistent, and the closed-loop demonstration is useful evidence of practical deployability. However, the significance is limited by the absence of parameter values and trial-level variance, by a motion model that is only tested on smooth yaw maneuvers, and by baselines that are not competitive by design. If the missing information is supplied and the robustness claims are appropriately scoped, the paper would be a useful contribution to cooperative AUV perception; in its current form, the quantitative claims are not fully supportable.","major_comments":[{"comment":"The key robustness claim for partial visibility rests on the constant-twist-acceleration random-walk prior in Eq. (16). With 1-3 visible LEDs, each pixel measurement provides at most 2m_k scalar constraints, so the unconstrained translational and rotational DOFs are supplied largely by the predicted pose. The experiments, however, only exercise a scripted 50 s yaw maneuver with smooth 30-degree rotations and hover segments; no step changes in acceleration, rapid translation, or high-bandwidth disturbance is included. Under an abrupt leader maneuver the prior can bias the few pixel residuals, and the visible LEDs cannot fully correct drift, so the reported mid-3-LED accuracy (Table II: e_t=0.036 m, e_r=0.042 rad) is evidence for the prior plus measurement combination only in the smooth-maneuver regime. Please add experiments with aggressive or step maneuvers, or explicitly restrict the claimed partial-visibility robustness to smooth relative motion.","section":"Section III-C (Eq. 16) and Section IV-A/E, Tables I-II"},{"comment":"All PSwPnP parameters are declared fixed for all experiments but never listed: Q_v, Q_a, R_proj, R_pix, R_det, lambda_clutter, p_survive, p_D, beta, sigma_pix, E_use, E_delete, E_confirm, and the initialization acceptance threshold in Eq. (8). Without these values the results cannot be reproduced, and it is impossible to assess whether the reported 0.032 m / 0.032 rad RMSE is robust or was tuned to the test conditions. Please provide a table of all parameter values and a sensitivity analysis, or at least an explicit statement of how the parameters were chosen on separate validation data.","section":"Section IV-A and Table I"},{"comment":"All quantitative comparisons are reported as single averages over 45 trials, with no standard deviations, confidence intervals, or statistical tests. Some improvements are large, but metrics such as P_{4->3}, P_{3->4}, ID agreement, and per-segment RMSEs can vary substantially across trials and target locations; the reader cannot judge whether differences are stable or dominated by a few trials. Please report trial-level variability (e.g., standard deviations, box plots, or per-location tables) and, where relevant, significance tests.","section":"Tables I and II (Section IV-D/E)"}],"minor_comments":[{"comment":"The frame-wise PnP baselines are not allowed any temporal filtering and freeze the previous pose when fewer than four LEDs are visible, so their poor performance under partial visibility is partly by construction; I recommend reframing these comparisons as sanity checks and foregrounding the w/o EKF ablation, which is the informative control for the value of the Lie-group predictor.","section":"Section IV-B and Table I"},{"comment":"The motion-capture settings in Fig. 5 list 'Shutter Speed: 90Hz'; this appears to be a frame rate rather than a shutter speed and should be corrected.","section":"Section IV-A, Fig. 5"},{"comment":"The 'real-time' feasibility claim would be strengthened by reporting the per-frame processing time and camera frame rate; currently no runtime numbers are given.","section":"Section V and Section IV-A"},{"comment":"The notation in Eq. (52) is confusing: H_{iℓ,k} is defined as ∂π/∂p_c times ∂p_c/∂δξ_ℓ, and then H_{iξ,k} is written as H_{iℓ,k} J_ℓ(ξ̂); please clarify which quantity is the measurement Jacobian and define δξ_ℓ explicitly before first use.","section":"Section III-F, Eq. (52)"},{"comment":"Please correct the typos: 'tranlation' in the Table I note, 'markder' in Section IV-C, 'intermdediate' in Section IV-A, and 'PIgment' in Fig. 2.","section":"Throughout"},{"comment":"References [20] and [21] are the same paper (Kim and Eustice, 'Real-time visual SLAM for autonomous underwater hull inspection using visual saliency') and should be merged into one entry.","section":"References"}],"recommendation":"major_revision","confidential_remarks":"The paper is likely a solid systems-and-application contribution after revision. The main risk to the headline claims is not the theory, which follows standard Lie-group EKF machinery, but the absence of parameter disclosure and variance information, and the mismatch between the constant-acceleration assumption and the benign test maneuvers. I would not reject the manuscript on these grounds; I would require the authors to supply the missing experimental detail and to temper or extend the partial-visibility robustness claims. The duplicate reference [20]/[21] and the scattered typos also suggest a final proofreading pass is needed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nI read AMR-Pose. The short version: it's a decent engineering contribution, and the experiments are more complete than most in this subfield, but the claims outrun the evidence in a few specific places. The core novelty is the integration, not any single component. The hardware—red central LED plus three blue LEDs with waterproof potting—is practical. The estimator is a standard Lie-group EKF on SE(3) with probabilistic marker association (Hungarian) and existence probabilities, and it switches to a pixel-level update when only 1–3 LEDs are visible. That switching formulation is new for AUVs, and the 4–3–4 visibility transition analysis is a nice piece of work. The ablation study is well designed and shows that each component (EKF, association, existence) contributes.\n\nWhat's good: motion capture ground truth over 45 trials, a visibility transition study, and a closed-loop demo. The math is internally consistent, following standard Lie-group EKF derivations. The quantitative gains over the frame-wise PnP baselines are large, and the qualitative plots show that the ablated variants break down in the expected way.\n\nWhere it's soft: the evaluation lacks statistical rigor. No error bars, no significance tests, no code or data. Dozens of parameters (Q_v, Q_a, R_proj, R_pix, etc.) are not reported, which makes the results hard to reproduce or assess for overfitting. The baselines are weak: EPnP-LM and GMLPnP with deterministic matching and hold-last-pose when fewer than four LEDs are visible. A stronger baseline would be a standard EKF with PnP updates, to isolate the contribution of the switching pixel-level update. The stress-test point about the constant-acceleration motion model is fair: with 1–3 LEDs, the pixel-level update constrains only a few DOFs, and the rest come from the prediction. The experiments only exercise a scripted 30-degree yaw maneuver, so the robustness to abrupt acceleration or high-bandwidth disturbance is unvalidated. That doesn't kill the paper, but it should temper the claim of robust partial-visibility tracking.\n\nWho benefits: anyone working on marker-based underwater relative pose, AUV leader-follower, or practical vision-based localization in degraded water. It deserves a serious referee—the work is within-subfield important and the experiments are substantial. I'd send it to review, but the reviewers should ask for error bars, parameter disclosure, code/data, and a stronger temporal baseline. Also an experiment with sharper motion would go a long way.\n\nRecommendation: engage with it.","headline":"Solid engineering integration of active LED markers with a switching EKF, but the evaluation is thin on statistical rigor and the constant-acceleration prior is untested under abrupt maneuvers.","tokens_in":17050,"tokens_out":4523,"would_cite":false,"duration_ms":38084,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Active LED markers and a probabilistic switching PnP estimator keep two AUVs' relative pose accurate and smooth even when only some LEDs are visible.","keywords":["active LED markers","relative pose estimation","autonomous underwater vehicles","Perspective-n-Point","Lie group SE(3)","probabilistic data association","visibility-adaptive filtering","leader-follower control"],"falsifier":"Run the same water-tank protocol but replace the smooth 30-degree yaw maneuvers with abrupt, high-jerk heading steps or a sudden dive that begins inside the three-LED segment, and check the estimates against motion-capture ground truth. Sharp error spikes during those segments would show that the constant-acceleration prior, not the switching visibility mechanism, is doing the work, whereas uninterrupted low error would strengthen the paper's claim.","tokens_in":16100,"feed_emoji":"🤖","tokens_out":10118,"duration_ms":90374,"temperature":0.7,"pith_summary":"The paper is trying to establish that two cooperating underwater robots can keep an accurate, continuous six-degree-of-freedom estimate of their relative pose even when the active LED marker on the leader is partially hidden, by letting the estimator switch between two measurement modes. Ordinary frame-by-frame PnP pose solving fails in this setting because it needs all four LEDs matched and ignores motion continuity; the proposed PSwPnP estimator instead propagates the pose on the Lie group $SE(3)$, associates detections probabilistically, and uses a pixel-level update when only one to three LEDs are available. The authors support the claim with water-tank experiments: 0.032 m translation RMSE and 0.032 rad rotation RMSE over 45 motion-capture-validated trials, better than frame-wise PnP baselines by large margins, plus a closed-loop leader-follower demonstration. If the claim holds, cooperative underwater navigation and formation control no longer need a full unobstructed marker view at every frame.","feed_headline":"Underwater robot pairs keep their relative pose as LEDs vanish","feed_subtitle":"Switching between full-marker pose updates and partial-marker pixel updates holds translation error at 0.032 m.","key_machinery":"The mechanism is a probabilistic switching Perspective-n-Point (PnP) estimator built on three coupled parts. First, the relative pose lives on the Lie group $SE(3)$, with a state of pose, twist, and twist acceleration; a discrete-time constant-acceleration motion model propagates the pose and its covariance between frames. Second, probabilistic marker association scores each detection against four LED trackers using either the tracker's image-plane prediction filter or the rigid-body reprojection prior, then solves a one-to-one assignment with a linear-assignment algorithm, while an existence-aware visibility manager smooths each LED's survival evidence in logit space and decides which LEDs are reliable. Third, the measurement update is visibility-adaptive: pose-level loose coupling for four visible LEDs and pixel-level tight coupling for one to three, so partial observations still constrain the full 6-DoF state through the motion model. The switching is what carries the argument: it lets the filter trade geometric richness of PnP against temporal continuity of direct pixel measurements.","core_discovery":"The central claim is that relative pose estimation between an active-marker-carrying leader and a camera-carrying follower can be made resilient to marker occlusion and detection clutter by fusing a Lie-group motion prior with visibility-adaptive measurements inside one recursive estimator, PSwPnP. When all four LEDs are reliably observed, the estimator solves a PnP problem and uses the result as a pose-level measurement; when one to three LEDs are observed, it uses their raw pixel projections in a tight-coupling update; when none are observed, it keeps only the $SE(3)$ prediction. Marker identity is maintained by probabilistic association and per-LED existence probabilities, so LEDs that disappear and reappear keep their labels. The experiments report that this design reduces translation RMSE to 0.032 m and rotation RMSE to 0.032 rad across 45 trials, while frame-wise PnP methods incur translation errors above 0.4 m and rotation errors near 0.9 rad under the same conditions.","pith_inferences":["The switching structure is not specific to four LEDs or to water; any sparse active or passive marker set with partial occlusion could use the same pose-level/pixel-level switching, provided a motion model for the target exists.","The experiments cover smooth yaw maneuvers in clear tank water; open-water turbidity, variable lighting, and abrupt leader accelerations remain untested, and those are the conditions most likely to break the brightness- and color-based LED segmentation.","When no LEDs are visible the filter keeps only the motion prediction, so the estimator degrades to dead reckoning; the paper does not quantify how long that prediction remains usable, which is a natural extension.","A direct observability analysis of one- and two-LED pixel updates could tell practitioners when the tight-coupling update is genuinely informative and when it only slows the drift of the motion prior."],"forward_implications":["A follower AUV can continue estimating the leader's six-degree-of-freedom pose while the leader turns and occludes two of its four LEDs, instead of freezing the last pose or dropping to dead reckoning.","Probabilistic association and logit-smoothed existence probabilities keep LED identities stable across disappearances and reappearances, so a recovered LED rejoins the estimate with its true label.","Temporal propagation on $SE(3)$, rather than any single-frame PnP solver, is the component that prevents pose jitter: removing it raises rotation RMSE from 0.032 rad to 0.587 rad in the reported trials.","The estimated pose is accurate and smooth enough to close a real-time leader-follower control loop, as demonstrated by the follower holding a roughly 1.8 m following distance while the leader translates and rotates."],"supporting_citations":[{"why":"Supplies the EPnP solver used for initialization and for the pose-level update when all four LEDs are visible.","marker":"[27]"},{"why":"Provides the uncertainty-aware GMLPnP baseline; the comparison shows frame-wise maximum-likelihood PnP cannot handle partial visibility.","marker":"[30]"},{"why":"Supplies the SE(3) exponential and log maps, adjoint, and left Jacobian used in Lie-group pose propagation.","marker":"[31]"},{"why":"Provides the manifold Gaussian-filter machinery for the rigid-body state prediction.","marker":"[32]"},{"why":"Gives the covariance-propagation formulas that carry pose uncertainty from the SE(3) prediction into the measurement updates.","marker":"[33]"},{"why":"Provides the linear-assignment solver used to enforce one-to-one detection-to-tracker matching.","marker":"[34]"},{"why":"Demonstrates active LED illumination as a viable underwater motion-measurement principle, motivating the marker design.","marker":"[22]"},{"why":"Supplies the AUV platforms used in the water-tank and closed-loop experiments.","marker":"[35]"}],"fun_headline_variants":["LED markers keep AUVs in sync even when some blink out","Switching between full and partial pose updates yields 0.032 m error","Probabilistic switching PnP stabilizes underwater robot relative pose","Partial LED views still give smooth AUV pose estimates","Underwater robot pairs hold pose accuracy as LEDs vanish"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The filter's continuity during partial visibility rests on the assumption that the leader's relative motion acceleration stays nearly constant between camera frames, with only small random disturbances; if the leader accelerates sharply while only one or two LEDs are visible, the predicted pose can drift and the few pixel measurements may not be enough to correct it.","fun_headline_variants_meta":{"raw":{"variants":["LED markers keep AUVs in sync even when some blink out","Switching between full and partial pose updates yields 0.032 m error","Probabilistic switching PnP stabilizes underwater robot relative pose","Partial LED views still give smooth AUV pose estimates","Underwater robot pairs hold pose accuracy as LEDs vanish"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000239,"raw_usage":{"total_tokens":1536,"prompt_tokens":988,"completion_tokens":548,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":604,"completion_tokens_details":{"reasoning_tokens":461}},"tokens_in":604,"tokens_out":548,"duration_ms":5659,"temperature":1.0,"reasoning_tokens":461,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-15T21:39:56.138739+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the same water-tank protocol but replace the smooth 30-degree yaw maneuvers with abrupt, high-jerk heading steps or a sudden dive that begins inside the three-LED segment, and check the estimates against motion-capture ground truth. Sharp error spikes during those segments would show that the constant-acceleration prior, not the switching visibility mechanism, is doing the work, whereas uninterrupted low error would strengthen the paper's claim.","supporting_citations":[{"cited_title":"The hungarian method for the assignment problem,","cited_arxiv_id":null,"evidence_quote":"Provides the linear-assignment solver used to enforce one-to-one detection-to-tracker matching."},{"cited_title":"Generalized maximum likelihood estimation for Perspective-n-point problem,","cited_arxiv_id":null,"evidence_quote":"Provides the uncertainty-aware GMLPnP baseline; the comparison shows frame-wise maximum-likelihood PnP cannot handle partial visibility."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"Provides the manifold Gaussian-filter machinery for the rigid-body state prediction."},{"cited_title":"Underwater space suit performance assessments part 1: Motion capture system development and validation,","cited_arxiv_id":null,"evidence_quote":"Demonstrates active LED illumination as a viable underwater motion-measurement principle, motivating the marker design."},{"cited_title":"A portable au- tonomous underwater vehicle with multi-thruster propulsion: Design, development, and vision-based tracking control,","cited_arxiv_id":null,"evidence_quote":"Supplies the AUV platforms used in the water-tank and closed-loop experiments."}],"review_version":1}