{"id":"4a3efc61-531e-4da8-8acf-d4db5b28d659","arxiv_id":"1908.08891","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":2,"one_line_summary":"A trinocular visual-inertial framework estimates a flexible, time-varying stereo baseline on a fixed-wing UAV by fusing IMU data, photometric alignment, and a calibrated wing model to produce depth maps.","lead":"This paper presents a visual-inertial system that estimates the changing relative positions of three cameras mounted on a fixed-wing drone's flexible wings and fuselage. The goal is to give fast-flying drones the long-range depth sensing of a wide stereo baseline without a heavy rigid sensor bar.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Wing-model prior may be calibrated and evaluated on the same flight; without a held-out quantitative pose comparison, the claimed baseline accuracy and generalization are unverified.","rationale":"The reader's weakest assumption and my load-bearing concern coincide: the paper does not demonstrate that the probabilistic wing model generalizes beyond the flight used to calibrate it. The framework is otherwise coherent: the EKF formulation is grounded in prior work, the runtime figures are plausible, and the qualitative reprojection and depth-map results support basic functionality. However, the absence of a held-out quantitative pose comparison is a conspicuous gap because the April-tag/side-camera apparatus already provides an independent reference at no additional hardware cost. This does not invalidate the approach, but it makes the current evidence insufficient for full acceptance. The appropriate verdict remains CONDITIONAL, pending a held-out validation experiment.","tokens_in":9365,"tokens_out":5532,"duration_ms":61959,"concrete_test":"Hold out a flight not used for wing-model calibration. On that flight, detect the April tags with the side cameras to obtain the reference transformation T_tagj_Cside,j (via Eq. 5) at each timestamp, and run the EKF with the pre-calibrated wing-model prior on the same data. Compare the EKF-estimated relative pose between wing and center camera to the tag-derived pose over the full flight, reporting median and 90th-percentile translation and rotation errors. If the errors remain within the few-cm/few-degree scatter observed in the calibration flight (Fig. 7), the prior generalizes; if errors exceed roughly 5 cm or 2 deg, the prior is overfit and the central claim is unsupported for new flights.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that the EKF estimates the non-rigid wing-to-center baseline accurately enough for long-range depth estimation. The evidence for this is qualitative: feature reprojection in Fig. 8 and sample depth maps in Fig. 9. No quantitative comparison of the estimated relative pose to an independent reference is reported. The only in-flight reference available, the April-tag/side-camera chain described in Sec. IV-C, is used to build the wing-model prior in Sec. VI-B. The paper never states that the evaluation flight is distinct from the calibration flight. If the prior is derived from the same flight used to demonstrate the system, the 'Calib-air' prior and the EKF result are partly fit to the deformation being estimated, so the claimed robustness across aerodynamic conditions is untested. This is load-bearing because Eq. (3) uses the wing model as a motion prior and the EKF fuses it as a Gaussian; a prior encoding one flight's specific deformation could bias relative pose estimates on a new flight, directly undermining long-range depth accuracy.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"This paper proposes a visual-inertial framework for estimating the time-varying relative pose between three cameras mounted on a fixed-wing UAV, where the outer cameras are on flexible wings and the center camera is in the fuselage. The relative pose between each wing camera and the center camera is estimated by an EKF that fuses relative IMU measurements, a photometric sparse image alignment term, and a probabilistic wing model learned from in-flight observations. The estimated poses are used to rectify image pairs and compute depth maps by block matching. The paper reports a wing-model calibration procedure using April tags and side cameras, hardware integration, real-world flight experiments, and runtime measurements. The central claim is that this non-rigid trinocular setup provides long-range depth estimation beyond what a rigid small-baseline rig can offer.","tokens_in":9516,"tokens_out":2386,"duration_ms":25117,"significance":"If the central claim holds, the work is a useful step for fixed-wing UAV perception, as it addresses a real deployment constraint: wide-baseline stereo on aeroelastically deforming wings. The paper's strengths are the complete system integration on a real platform, the hardware-synchronized multi-camera-IMU sensor design, the in-flight wing-model calibration procedure, and the demonstration that a photometric update with a good prior runs efficiently on the tested hardware. However, the current evidence is largely qualitative. The depth maps and pose estimates are shown as images, but no quantitative comparison against independent ground truth is provided, and the relationship between the calibration flight and the evaluation flight is not stated. These gaps directly affect the strength of the main claim that the estimated baseline is accurate enough for long-range depth estimation.","major_comments":[{"comment":"The core claim that the EKF accurately estimates the time-varying wing-to-center baseline is supported only qualitatively. The paper shows reprojected features and sample depth maps, but it does not report any quantitative error metric for the estimated relative pose TCj/Cc or for the depth maps. Given that the motivation is long-range depth accuracy, please add numbers: for example, root-mean-square or median errors of the estimated relative translation and rotation against an independent reference, and depth error or disparity error statistics when ground-truth or a held-out reference is available.","section":"Sec. VI-C and VI-D, Figs. 8 and 9"},{"comment":"The manuscript does not state whether the flight used to build the probabilistic wing model is the same flight used for the qualitative demonstrations in Figs. 8 and 9. If the 'Calib-air' prior is derived from the same flight on which the system is then evaluated, the EKF result is partly fit to the deformation of that flight, and the claimed robustness across aerodynamic conditions is untested. Please state explicitly whether the calibration and evaluation flights are distinct, and ideally evaluate on a held-out flight or compare the EKF estimate frame-by-frame against the independent April-tag/side-camera reference from Sec. IV-C.","section":"Sec. IV-C and Sec. VI-B"},{"comment":"The wing-model prior appears both as the motion prior in the photometric objective (Eq. (3)) and as the Gaussian fused in the EKF, but the manuscript does not specify how the prior covariance is used or how its weight is set relative to the vision update. Since a too-confident prior would dominate the visual-inertial measurements and effectively reproduce the calibration fit, please clarify the fusion formulation, report the covariance values used, and provide a sensitivity analysis to the prior weight.","section":"Eq. (3) and Sec. IV-A"}],"minor_comments":[{"comment":"The runtime in Table I is measured on an Intel i7-4800MQ, while the onboard computer is an UP Squared with an Intel Atom at 1.6 GHz; please clarify whether the stated 'some margin' conclusion applies to the actual onboard platform or only to the more powerful comparison machine.","section":"Table I"},{"comment":"The state vector in Eq. (1) includes IMU biases for both cameras, but the propagation and update equations for these biases are not given in the paper; citing [19] is acceptable, but a brief description of how the biases are modeled would improve readability.","section":"Sec. IV-A"},{"comment":"The depth maps in Fig. 9 are single-shot examples without a color scale or quantitative depth legend, making it hard for the reader to judge the actual depth range or to compare the proposed method against the two priors beyond visual inspection.","section":"Fig. 9"},{"comment":"There are a few typographical and formatting issues, such as the text running into figure captions in Sec. VI-B ('The observed 150 200 250 300 3500.26'), which should be cleaned up in the final version.","section":"General"}],"recommendation":"major_revision","confidential_remarks":"The main risk is the calibration/evaluation circularity: the wing model is fitted to in-flight data and may be evaluated on the same flight. The authors likely have additional data, but the paper as written does not demonstrate generalization. If they can provide quantitative pose or depth evaluation against an independent reference and clarify the flight separation, the work would be a solid IROS contribution. The novelty relative to the authors' prior Flexible Stereo work [1] is real but incremental; the trinocular setup and the in-flight calibration procedure are the main new elements."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Colleague,\n\nThis is a genuine but incremental step beyond the authors' own Flexible Stereo work. What's new: a trinocular fixed-wing setup (center plus left/right wing-tip cameras), an EKF that estimates the time-varying baseline while compensating IMU biases, a photometric sparse-image-alignment update, and real flight data—the first real-world validation of this flexible-baseline idea. The hardware design and wing-model calibration using April tags observed from side cameras is a sensible, practical procedure, and the runtime table shows the pipeline can run at 10 Hz on an Atom-class board. The paper is clear about its scope: it assumes the center camera pose is known from another source and does not claim full map estimation.\n\nThe main soft spot is exactly what the stress-test flags: the probabilistic wing model is built from in-flight April-tag observations, and the paper never states that the evaluation flights are different from the calibration flight. If the same flight's data produced both the 'Calib-air' prior and the depth maps that demonstrate the system, then the prior is partly fit to the deformation being estimated, and the qualitative improvement in Figs. 8–9 is far less convincing. This is a real, addressable concern; the fix is easy—report on a held-out flight or at least quantify the wing-model error on a different day.\n\nA second soft spot is the absence of quantitative accuracy evaluation. The relative pose between wing and center cameras could be checked against the same April-tag chain used for calibration, but instead we get feature reprojection images and sample depth maps. No RMSE, no error bars, no depth ground truth. No code or data is released, so independent reproduction is not possible. These gaps leave the central accuracy claim unverified, but they do not sink the core idea. The EKF formulation is inherited from peer-reviewed work, the inertial and photometric updates are independent of the wing model, and the engineering is credible. The citation pattern is unremarkable; the authors build on their own prior paper and cite the relevant related work.\n\nThis is useful reading for anyone working on wide-baseline stereo on deformable platforms. I would send it to peer review with the expectation of a major revision, because the problem is important for fast fixed-wing navigation and the gaps are fixable with a moderate amount of additional experiments. I'd also bring it to a reading group if UAV perception is on the agenda.","headline":"Genuine incremental step with real flight tests, but held-out evaluation and quantitative pose/depth error are missing; the central accuracy claim is not yet proven.","tokens_in":10041,"tokens_out":3612,"would_cite":true,"duration_ms":33791,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper aims to show that a fast fixed-wing UAV can estimate its own time-varying trinocular baseline well enough to produce long-range dense depth maps for navigation and mapping.","keywords":["visual-inertial odometry","non-rigid stereo","wide-baseline depth estimation","fixed-wing UAV","extended Kalman filter","photometric image alignment","dense reconstruction","wing deformation model"],"falsifier":"Fly the platform through a maneuver envelope not represented in the calibration flight, while independently measuring the wing-camera poses from the side cameras' fiducial markers; if the marker-derived poses systematically drift outside the EKF's predicted uncertainty as airspeed or gust level increases, the fixed wing-model prior is not generalizing and the central claim collapses.","tokens_in":9147,"feed_emoji":"🛩️","tokens_out":9865,"duration_ms":88795,"temperature":0.7,"pith_summary":"This paper tries to establish that a fast fixed-wing UAV does not need a rigid wide stereo mount to get long-range depth: the flexible wing itself can serve as a wide baseline, as long as the camera at each wing tip and the camera in the fuselage are continuously re-aligned. The alignment is done by fusing relative inertial measurements from the wing-tip and center IMUs, a photometric alignment of the overlapping images, and a probabilistic wing-deformation model in an extended Kalman filter. The result is time-varying camera poses that support dense depth maps from both the full wing-to-wing baseline and the shorter wing-to-center baselines, which can be registered into a map for local replanning. A sympathetic reader should care because a rigid stereo pair inside a small fuselage is fundamentally limited in range, while off-the-shelf long-range sensors are heavy or expensive; this approach claims to get the benefit of a wide baseline from low-cost visual-inertial hardware.","feed_headline":"Wing flex becomes the stereo baseline for fast UAV depth","feed_subtitle":"An EKF fuses wing-mounted IMUs, photometric tracking, and a wing model to stretch stereo depth toward long range.","key_machinery":"The load-bearing object is a relative extended Kalman filter for each wing-center camera pair, with a state that includes the relative rotation quaternion, translation, angular velocities, linear accelerations, and IMU biases. Its job is to stay close to the true wing-to-center transform between frames: the relative IMU propagation gives the prior, the photometric refinement corrects it with image intensities, and the probabilistic wing model—a Gaussian mean and covariance for the same transform learned from in-flight fiducial-marker observations—pulls the estimate back when vision fails. The photometric step is what makes the translation full-scale rather than up-to-scale, which is the property that turns a flexible baseline into metric depth.","core_discovery":"The discovery is that the time-varying relative pose of cameras mounted on a flexible fixed-wing aircraft can be estimated tightly enough to generate dense depth maps at ranges unreachable by a rigid small-baseline rig, and that this can be done in real time with low-cost sensors. The estimator is an EKF in relative form whose state covers, for each wing camera, the rotation and translation to the center camera, angular velocities, accelerations, and IMU biases. The EKF propagates the baseline with IMU data, then a photometric sparse image alignment—using the predicted pose to project center-camera feature patches into the wing image and minimizing the intensity difference with a constrained Gauss-Newton solver—provides a full-scale pose update. That update is fused with a Gaussian wing model, a mean and covariance of the wing-to-center transform calibrated in flight via fiducial markers observed by side cameras. With the corrected poses, stereo rectification and block matching produce depth maps from the full baseline and from the half baselines, and the center camera's rigid link to the autopilot lets those maps be geo-referenced.","pith_inferences":["A clear next experiment is to split calibration and evaluation flights: learn the wing-model Gaussian on one flight, then fly on a different day or airspeed regime and check whether EKF poses stay inside the filter's predicted uncertainty; this would settle whether the fixed prior generalizes.","Conditioning the wing model on airspeed or measured load factor should remove the largest expected bias, since the paper's own take-off data show the relative transform shifting abruptly and the future-work notes already point toward a cantilever-beam model with airspeed as an input.","The modular filter structure invites adding extra camera-IMU pairs, such as on the tail or nose, to widen the field of view or add short-range baselines without re-deriving the estimator; each pair is just another relative EKF."],"forward_implications":["A fixed-wing UAV with a wide flexible baseline can produce depth maps at ranges where a rigid in-fuselage stereo rig would have one-pixel disparity limits, making long-range navigation feasible with low-cost sensors.","The same three cameras give two half-baseline pairs for near-field tasks such as landing and obstacle avoidance, and one full-baseline pair for distant terrain, without extra hardware.","Because the center camera is rigidly attached to the fuselage and autopilot, the depth maps can be transformed directly into a geo-referenced map for local replanning, rather than serving only reactive avoidance.","The wing-model calibration using side cameras and fiducial markers is a one-time procedure per UAV type, so the operational sensor suite remains just the three cameras and IMUs.","The modular EKF and per-frame runtime mean the pipeline can run at 10 Hz on small onboard computers and can accommodate extra camera-IMU rigs to widen the field of view."],"supporting_citations":[{"why":"Introduces the earlier two-camera flexible stereo formulation and the probabilistic wing-model constraint that this three-camera system extends to real flight.","marker":"[1]"},{"why":"Supplies the relative EKF formulation that propagates and fuses IMU and vision measurements between two rigid bodies.","marker":"[19]"},{"why":"Provides the sparse photometric image-alignment approach used to refine the wing-to-center relative pose.","marker":"[20]"},{"why":"Supplies the constrained Gauss-Newton solver and pyramidal optimization used in the photometric refinement step.","marker":"[23]"},{"why":"Provides the rectification algorithm that turns the estimated pose into aligned stereo image pairs for dense matching.","marker":"[24]"},{"why":"Calibrates the rigid camera-to-IMU transformations that anchor the wing-model transformation chain.","marker":"[25]"},{"why":"Implements the stereo block-matching routine used to produce the dense depth maps from the rectified images.","marker":"[29]"}],"fun_headline_variants":["Flexible trinocular rig extends UAV depth range","Wing-mounted cameras yield long-range depth in flight","Non-rigid trinocular EKF boosts drone perception","Adaptive baseline from wing flex for aerial mapping","Real-time depth from a bending drone camera rig"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The whole estimate leans on a fixed probabilistic wing model learned from one in-flight calibration, and if that prior does not match the deformation on the flight being evaluated, the baseline estimates will be biased.","fun_headline_variants_meta":{"raw":{"variants":["Flexible trinocular rig extends UAV depth range","Wing-mounted cameras yield long-range depth in flight","Non-rigid trinocular EKF boosts drone perception","Adaptive baseline from wing flex for aerial mapping","Real-time depth from a bending drone camera rig"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000113,"raw_usage":{"total_tokens":1047,"prompt_tokens":909,"completion_tokens":138,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":525,"completion_tokens_details":{"reasoning_tokens":62}},"tokens_in":525,"tokens_out":138,"duration_ms":2292,"temperature":1.0,"reasoning_tokens":62,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T11:25:52.277086+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Fly the platform through a maneuver envelope not represented in the calibration flight, while independently measuring the wing-camera poses from the side cameras' fiducial markers; if the marker-derived poses systematically drift outside the EKF's predicted uncertainty as airspeed or gust level increases, the fixed wing-model prior is not generalizing and the central claim collapses.","supporting_citations":[{"cited_title":"Flexible stereo: Con- strained, non-rigid, wide-baseline stereo vision for ﬁxed-wing aerial platforms,","cited_arxiv_id":null,"evidence_quote":"Introduces the earlier two-camera flexible stereo formulation and the probabilistic wing-model constraint that this three-camera system extends to real flight."},{"cited_title":"Col- laborative Stereo,","cited_arxiv_id":null,"evidence_quote":"Supplies the relative EKF formulation that propagates and fuses IMU and vision measurements between two rigid bodies."},{"cited_title":"SVO: semidirect visual odometry for monocular and multicamera systems,","cited_arxiv_id":null,"evidence_quote":"Provides the sparse photometric image-alignment approach used to refine the wing-to-center relative pose."},{"cited_title":"A compact algorithm for rectiﬁcation of stereo pairs,","cited_arxiv_id":null,"evidence_quote":"Provides the rectification algorithm that turns the estimated pose into aligned stereo image pairs for dense matching."},{"cited_title":"Extending kalibr: Calibrating the extrinsics of multiple IMUs and of individual axes,","cited_arxiv_id":null,"evidence_quote":"Calibrates the rigid camera-to-IMU transformations that anchor the wing-model transformation chain."},{"cited_title":"The OpenCV Library,","cited_arxiv_id":null,"evidence_quote":"Implements the stereo block-matching routine used to produce the dense depth maps from the rectified images."}],"review_version":1}