{"id":"daae7fde-1284-438c-a80d-79aceaf4cf55","arxiv_id":"2606.05915","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":6.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"CamFlow+ is a hybrid-basis framework for 2D camera motion estimation that relaxes single-plane constraints using physical, stochastic, and depth-derived bases, with evaluation on a masked camera-motion benchmark and stabilization applications.","lead":"CamFlow+ introduces a hybrid motion bases framework for 2D camera motion estimation in dense flow space, combining homography-derived, stochastic, and depth-translational bases to handle translation and parallax. This could improve video stabilization and motion analysis in scenes with depth variation beyond traditional planar assumptions.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"No significant objection identified","rationale":"Reader correctly flagged the hybrid-regularization assumption as central and noted the abstract-only limitation. With full text now referenced, the same assumption remains the only plausible soft spot, but no concrete flaw is detectable without the missing implementation or ablation details. Verdict therefore stays UNVERDICTED pending those details.","tokens_in":1744,"tokens_out":251,"duration_ms":13937,"concrete_test":"Re-run the GHOF-Cam sparse/dense error tables after replacing the supplied depth map with one perturbed by additive Gaussian noise at σ = 5 % of median depth; if the reported improvement over baselines disappears, the depth-translational bases are the dominant untested assumption.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The abstract describes a hybrid basis construction (homography physical + stochastic + depth-translational) plus depth-aware smoothness that directly targets the known failure modes of single-plane homography. The evaluation isolates camera motion via GHOF-Cam masking and reports both quantitative gains and a user-study preference. No internal contradiction or unsupported leap is visible from the provided description; the weakest_assumption is plausible given the stated components.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces CamFlow+, a hybrid-basis framework for 2D camera motion estimation in dense flow space. It combines homography-derived physical bases, stochastic bases sampled from homography flows, and depth-translational bases derived from depth and intrinsics, plus a depth-aware smoothness term, to relax the single-plane assumption while preserving regularity. Evaluation on the GHOF-Cam benchmark (which masks dynamic objects and occlusions from an optical-flow dataset) reports improvements in sparse and dense camera-motion estimation; in video stabilization it improves global and local stability and achieves the highest top-1 preference in a blind user study.","tokens_in":1800,"tokens_out":559,"duration_ms":17420,"significance":"If the quantitative gains and user-study results hold under rigorous controls, the hybrid construction offers a principled way to capture translation-induced parallax without piecewise-planar or mesh-based overhead, which could benefit stabilization, SfM, and video processing pipelines. The explicit construction of bases from homography and depth is a strength that keeps the model interpretable.","major_comments":[{"comment":"Abstract and §4 (Experiments): the claim of improvement on GHOF-Cam is stated without any reported error metrics, baseline comparisons, ablation tables, or statistical significance tests; because the central claim is empirical superiority, the absence of these numbers in the provided text prevents verification that the hybrid bases actually outperform existing methods by a meaningful margin.","section":"Abstract, §4"},{"comment":"§3.3 (Depth-aware smoothness): the term is described as regularizing translation-induced parallax in continuous-depth regions while preserving boundaries, but no derivation or weighting schedule is supplied; without an equation or sensitivity analysis it is unclear whether the term is load-bearing or could be replaced by a standard smoothness prior.","section":"§3.3"}],"minor_comments":[{"comment":"The GHOF-Cam construction (masking procedure, exact benchmark split) should be detailed in a dedicated subsection or supplementary material so that the isolation of camera motion can be reproduced.","section":"§4.1"},{"comment":"Notation for the three basis families (physical, stochastic, depth-translational) is introduced in the abstract but not consistently labeled in the method section; a single table summarizing their definitions and dimensions would improve clarity.","section":"§3"},{"comment":"The statement that code and datasets will be released is welcome; the camera-ready version should include the exact commit hash or DOI once available.","section":"Abstract"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for the constructive feedback on the empirical presentation and the depth-aware smoothness term. We address each major comment below and will update the manuscript to improve clarity and verifiability.","responses":[{"response":"The experiments section (§4) contains tables reporting quantitative error metrics (e.g., endpoint error for sparse and dense camera-motion estimation), direct comparisons against baselines, and ablation studies on the hybrid bases. However, the abstract summarizes results qualitatively and §4 does not explicitly highlight statistical significance. We will revise the abstract to include key numerical gains and augment §4 with explicit statistical tests or confidence intervals to allow direct verification of the claimed improvements.","revision_made":"yes","referee_comment":"[Abstract, §4] Abstract and §4 (Experiments): the claim of improvement on GHOF-Cam is stated without any reported error metrics, baseline comparisons, ablation tables, or statistical significance tests; because the central claim is empirical superiority, the absence of these numbers in the provided text prevents verification that the hybrid bases actually outperform existing methods by a meaningful margin."},{"response":"We agree that §3.3 currently provides only a high-level description. In the revised manuscript we will add the explicit equation for the depth-aware smoothness term, derive it from the depth map and translation-induced flow, specify the weighting schedule used during optimization, and include a sensitivity/ablation study comparing it against a standard smoothness prior to demonstrate its contribution.","revision_made":"yes","referee_comment":"[§3.3] §3.3 (Depth-aware smoothness): the term is described as regularizing translation-induced parallax in continuous-depth regions while preserving boundaries, but no derivation or weighting schedule is supplied; without an equation or sensitivity analysis it is unclear whether the term is load-bearing or could be replaced by a standard smoothness prior."}],"tokens_in":1410,"tokens_out":404,"duration_ms":19184,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The main point is that this paper builds CamFlow+ as a hybrid set of bases for 2D camera motion in dense flow space. It takes homography-derived physical flows, adds stochastic samples from those flows, and includes depth-translational bases from depth maps and intrinsics, then adds a depth-aware smoothness term to handle parallax without breaking at depth edges.\n\nWhat stands out is the direct attack on the single-plane limit that breaks standard homography when the camera translates. The GHOF-Cam benchmark, made by masking dynamic objects and bad regions from an optical flow set, is a reasonable way to isolate the camera component. The stabilization section reports better global and local stability plus the top spot in a blind user preference test.\n\nThe soft spot is the complete lack of numbers. The abstract gives no error rates, no baseline comparisons, no ablation on which base matters most, and no sense of how much the depth term helps versus hurts. Without those, it is hard to know if the hybrid actually moves the needle or just adds complexity. The full paper may fix this, but the current description leaves the strength untested.\n\nThis is for people who work on video stabilization or structure-from-motion pipelines that need to cope with real depth variation. A reader who already knows the homography failure modes will see the motivation clearly. It deserves a serious referee because the problem is real, the components are spelled out, and the user study adds a practical check, even if the quantitative case still needs to be made.","headline":"CamFlow+ mixes homography, stochastic, and depth bases in flow space to relax planar assumptions for camera motion, with stabilization wins in a user study but no numbers visible yet.","tokens_in":2274,"tokens_out":388,"would_cite":false,"duration_ms":20682,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"CamFlow+ uses hybrid motion bases to estimate 2D camera motion in dense flow space, handling translation and parallax without single-plane assumptions.","keywords":["camera motion estimation","video stabilization","hybrid motion bases","dense optical flow","depth-aware regularization","homography","parallax handling"],"falsifier":"On a video sequence with complex non-planar depth variations and parallax where CamFlow+ either fails to outperform homography methods or produces visible artifacts in the stabilized output.","tokens_in":2649,"feed_emoji":"📹","tokens_out":602,"duration_ms":20840,"temperature":0.7,"pith_summary":"Existing methods for 2D camera motion estimation rely on homography or mesh models that assume planar scenes or piecewise planarity, limiting their ability to handle camera translation, depth variation, and local parallax. CamFlow+ introduces a hybrid framework in dense-flow space that combines homography-derived physical bases, stochastic bases from homography flows, and depth-translational bases from depth and intrinsics. A depth-aware smoothness term regularizes parallax in continuous depth areas while keeping changes at boundaries. This leads to better sparse and dense motion estimates on a dedicated benchmark and improved global and local stability in video stabilization, with top preference in user studies.","feed_headline":"Hybrid bases improve camera motion estimates in flow space","feed_subtitle":"Combining physical, stochastic and depth bases relaxes planar assumptions for better 2D motion and stabilization results.","key_machinery":"The hybrid-basis framework that mixes physical, stochastic, and depth-translational bases in dense-flow space along with depth-aware smoothness regularization.","core_discovery":"CamFlow+ represents 2D camera motion directly in dense-flow space by combining homography-derived physical bases, stochastic bases sampled from homography flows, and depth-translational bases derived from depth and camera intrinsics, while adding a depth-aware smoothness term to regularize translation-induced parallax.","pith_inferences":["If the hybrid bases generalize, they could reduce the need for scene-specific tuning in stabilization pipelines.","Extending the depth-translational bases to dynamic scenes might allow better separation of camera and object motion.","The dense-flow representation could integrate with learning-based methods for end-to-end training."],"forward_implications":["CamFlow+ improves accuracy in sparse and dense camera-motion estimation on the GHOF-Cam benchmark.","It enhances global and local stability in digital video stabilization applications.","It achieves the highest top-1 preference rate in a blind user study for stabilized videos.","The approach relaxes the single-plane constraint while preserving camera-motion regularity."],"fun_headline_variants":["CamFlow+ hybrid bases represent camera motion in flow space","Physical stochastic and depth bases combined for camera motion","Hybrid bases relax single plane constraint with depth term","CamFlow+ mixes homography and depth bases in dense flow"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The combination of the three base types and the depth-aware term captures real-world camera translation and parallax without introducing artifacts or needing per-scene adjustments.","fun_headline_variants_meta":{"raw":{"variants":["CamFlow+ hybrid bases represent camera motion in flow space","Physical stochastic and depth bases combined for camera motion","Hybrid bases relax single plane constraint with depth term","CamFlow+ mixes homography and depth bases in dense flow"]},"model":"grok-4.3","cost_usd":0.008301,"raw_usage":{"total_tokens":3755,"prompt_tokens":654,"num_sources_used":0,"completion_tokens":62,"cost_in_usd_ticks":83012000,"prompt_tokens_details":{"text_tokens":654,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":3039,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":654,"tokens_out":62,"duration_ms":22513,"temperature":1.0,"reasoning_tokens":3039,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-28T01:56:34.239024+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"On a video sequence with complex non-planar depth variations and parallax where CamFlow+ either fails to outperform homography methods or produces visible artifacts in the stabilized output.","supporting_citations":[],"review_version":1}