{"id":"79f93887-d9fa-4481-8240-0a3e535acab1","arxiv_id":"2608.00950","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"Swimm3R couples a scattering-aware, feed-forward structure-from-motion backbone with underwater Beta splatting to reconstruct and render 3D scenes from turbid underwater video, improving rendering PSNR and localization over prior underwater splatting methods.","lead":"The paper introduces Swimm3R, a system that estimates underwater camera motion and 3D structure directly from hazy, color-shifted video, then renders maps with a new scattering-aware splatting method. It reports cleaner seafloor geometry and a 1.47 dB PSNR gain over a leading underwater renderer, plus a small boost in downstream localization accuracy.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Barbados PSNR gain may be inflated by test-time pose refinement; a no-refinement check is needed.","rationale":"The reader identified test-time pose refinement as the weakest assumption, and I agree this is the most load-bearing concern for the PSNR claim. However, I weight it slightly differently: the localization claim in Tab. V is not affected by pose refinement, and the FLSea results provide an independent, externally referenced validation of the SfM improvements. Thus the central contribution is not wholly invalidated by the Barbados protocol. The concern is real but addressable with a concrete evaluation change, so the verdict remains CONDITIONAL, consistent with the reader's assessment. I partially agree with the reader's weakest_assumption: the pose-refinement issue is the primary soft spot, while the GLUEMAP-reference concern is secondary because all methods share the same reference and the FLSea pose metrics reinforce the trend.","tokens_in":13891,"tokens_out":8765,"duration_ms":97051,"concrete_test":"Re-run the Table IV comparison with test-time pose refinement disabled. Fix each test camera pose at the coarse pose predicted by the feed-forward SfM backbone (or at the nearest-neighbor training pose), and compute PSNR/SSIM/LPIPS for all renderers without any pose optimization. If UWBS's average PSNR advantage over WaterSplatting drops below ~0.5 dB or changes sign, the reported 1.47 dB rendering gain is substantially an artifact of pose overfitting. Report both the no-refinement and refined numbers to quantify the effect.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The headline rendering claim (Tab. IV: UWBS improves average PSNR by 1.47 dB over WaterSplatting on the Barbados dataset) is evaluated under the Sec. IV-E protocol: each held-out test camera pose is refined for 300 steps against its own test image while the map and medium field are fixed. This measures the renderer's ability to fit a pose to the image, not pure novel-view synthesis. Since Beta kernels (the primary difference between UWBS and WaterSplatting) may respond differently to pose optimization, the relative PSNR ordering could shrink or change if pose refinement were removed. The localization claim (Tab. V) is not affected because it does not use test-time pose refinement, and the FLSea depth/pose results (Tab. II) provide an external anchor. However, the Barbados dataset is unreleased, and the paper itself states no metric ground-truth poses are available (Sec. V, Limitations), so the GLUEMAP reference used for RRA/RTA is an additional assumption. The weakest load-bearing link is therefore the PSNR comparison on an unreleased dataset under a protocol that permits pose overfitting.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes Swimm3R, a unified underwater 3D reconstruction framework that combines a LoRA-fine-tuned MASt3R-SfM backbone with an underwater physics head (predicting attenuation, backscatter, and veiling light) and an Underwater Beta Splatting (UWBS) renderer using Beta kernels, an explicit medium field, and scattering-aware gradient preconditioning. A new four-scene Barbados GoPro dataset is introduced. On the FLSea test sequences, Swimm3R reports improved depth (AbsRel 0.2132 vs. 0.2247 for MASt3R-SfM) and pose accuracy; on Barbados, it reports denser point clouds, better test-frame rendering PSNR (1.47 dB over WaterSplatting under Swimm3R initialization), and improved downstream localization (RRA@15 and RTA@15 gains of 2.0 and 2.4 points). The central claim is that medium-aware SfM plus Beta-based splatting improves underwater geometry, rendering, and localization in scattering conditions.","tokens_in":14184,"tokens_out":6121,"duration_ms":67519,"significance":"If the results hold, the paper addresses a real limitation of underwater 3D reconstruction: the fragility of SfM poses and sparse points under scattering and attenuation. Distilling in-air geometric priors through a physics-based image-formation head is a timely idea, and the external FLSea evaluation provides a non-circular anchor for the SfM component. The explicit underwater medium field, Beta kernels, and SGG preconditioner are inventive, and the ablations are internally consistent. The paper is also admirably candid about the absence of metric ground-truth poses on Barbados and about the comparative nature of the pose numbers. However, the headline Barbados rendering result is measured under a test-time pose-refinement protocol, and the SfM training objective is underspecified with respect to the depth variable, so the empirical support is not yet fully load-bearing.","major_comments":[{"comment":"The reported 1.47 dB PSNR gain over WaterSplatting is obtained after refining each test camera pose for 300 steps against its own held-out image while the map and medium field are fixed. This measures pose-fitting plus rendering, not pure novel-view synthesis, and can differentially favor a renderer whose Beta kernels interact favorably with pose optimization. Since this is the headline rendering claim on a new dataset, please also report PSNR/SSIM/LPIPS without test-pose refinement (or with a fixed common pose), and discuss whether the relative ordering persists. The localization claim in Tab. V is not affected by this issue, but the rendering claim is currently tied to this protocol.","section":"Sec. IV-E, Tab. IV"},{"comment":"The physics-head forward and inverse losses require a metric ray depth z_n (Eq. 2), but the manuscript never specifies where z_n comes from during SfM training. If z_n is the student's own predicted depth, then the inverse restoration in Eq. (4) and the forward loss in Eq. (5) are trained with self-generated depth, which is a circularity concern and should be stated. If z_n is obtained from the teacher pointmap, stereo, known geometry, or another source, that source and the gradient flow need to be described. This is load-bearing for interpreting the SfM ablations in Tab. VI.","section":"Sec. III-B, Eqs. (2)-(6)"},{"comment":"The Barbados 'SfM map quality' metrics (epipolar error and 1-px inlier ratio) are computed under each method's own poses, as stated in Sec. IV-B. They are internal-consistency measures, not accuracy relative to a shared reference. With no metric ground-truth geometry on Barbados, the conclusion that Swimm3R yields 'more continuous and geometrically coherent scene structures' is supported only by self-consistent numbers and qualitative figures. Please make this caveat prominent in Sec. IV-D and in the Tab. III caption, and avoid implying that lower self-epipolar error directly translates to metric accuracy.","section":"Sec. IV-B, Tab. III"},{"comment":"On Barbados, GLUEMAP poses serve as the reference for RRA@15 and RTA@15. The paper acknowledges in Sec. V that no metric ground-truth poses are available, but the main-text sentence in Sec. IV-B and the Tab. V caption should explicitly state that all Barbados pose numbers are relative to GLUEMAP and therefore comparative rather than absolute. This is not fatal, but it should be made equally prominent at the point of use, not only in the limitations paragraph.","section":"Sec. IV-B, Tab. V, Sec. V"}],"minor_comments":[{"comment":"The ambient convolutional field A_phi is described as a 1x1 CNN with ReLU hidden activations, but its input/output dimensionalities and the SH degree used are not specified. Please clarify the architecture and how B_inf, beta_B, and beta_D are split from its output.","section":"Sec. III-C, Eq. (10)"},{"comment":"The notation [tilde J]_0^1 for clipping to [0,1] is used without definition. Please define it explicitly in the text or with a comment after the equation.","section":"Eq. (4)"},{"comment":"The caption says 'test-frame rendering quality' but the protocol includes 300-step per-test-pose refinement. Please state 'with test-pose refinement' in the caption or table title so that readers are not misled into interpreting the numbers as pure held-out novel-view synthesis.","section":"Tab. IV caption"},{"comment":"The FLSea pose evaluation uses COLMAP on clean sequences as reference. Since Dark3R and MASt3R-SfM are evaluated against the same reference, the comparison is fair; however, a sentence clarifying that FLSea pose numbers are scale-free and reference-relative would help.","section":"Sec. IV-B"},{"comment":"The paper establishes a new dataset but no release URL or availability statement is provided. If the dataset is to be a contribution, please include a link or state the intended release plan in the paper.","section":"Barbados dataset"}],"recommendation":"major_revision","confidential_remarks":"The Barbados dataset is not yet released, and the headline rendering comparison depends on a test-time pose-refinement protocol. I would like to see either code/data availability or at least the no-refinement rendering results before endorsing the quantitative claims. The FLSea results and the honest limitation statement are encouraging, so this is a revision rather than a rejection."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Worth a serious read and a serious referee. The core idea is genuinely new: adapt MASt3R-SfM to underwater imagery by fine-tuning with LoRA and adding a physics head trained with forward and inverse losses on paired clean-underwater data. The FLSea depth/pose results support it, beating MASt3R-SfM and Dark3R, and the ablations show each loss component matters. The UWBS renderer also has substance: Beta kernels plus an explicit medium field with scattering-aware gradients, and the localization experiment cleanly isolates the renderer by fixing the SfM map. The authors are honest in the Limitations section that there is no metric ground truth on the new Barbados sequences.\n\nNow the soft spots, in order of weight. First, the headline PSNR gain (1.47 dB over WaterSplatting) is measured after 300 steps of test-time pose refinement per test frame against the test image itself. That is pose fitting, not pure novel-view synthesis, and it may favor the flexible Beta kernels differently than it favors the Gaussian baseline. The protocol is applied to all methods, so it is a fair comparison, but the absolute gap could shrink if refinement were removed. A no-refinement ablation would settle it quickly. Second, the localization reference on Barbados is GLUEMAP, not ground truth, so those numbers are comparative rather than exact, as the authors concede. Third, no code or data release; the Barbados dataset is not public, so independent replication is currently impossible.\n\nNone of this is fatal. The FLSea results are an external anchor, the ablations are internally consistent, and the physics head is not purely fitting the evaluation metric (forward and inverse losses are supervised on paired data). But the magnitude of the Barbados gains is unverifiable as presented. I would condition acceptance on a no-refinement Barbados PSNR table and on releasing data and code. If those hold, this is an advance the underwater 3D community will actually use.","headline":"Credible, well-ablated adaptation of feed-forward SfM to underwater imaging; the headline PSNR gains rest on an unreleased dataset with test-time pose refinement, so treat the magnitudes with caution.","tokens_in":14718,"tokens_out":3096,"would_cite":true,"duration_ms":34063,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"Swimm3R claims that scattering-degraded underwater video can be reconstructed into dense, restored 3D maps by distilling in-air geometric priors into a medium-aware SfM backbone and rendering with Beta splatting.","keywords":["underwater 3D reconstruction","structure-from-motion","Gaussian splatting","Beta splatting","scattering medium","visual localization","image restoration","underwater dataset"],"falsifier":"Run Swimm3R, MASt3R-SfM, and WaterSplatting on an underwater scene with known metric poses and depth—for example a tank with calibrated cameras and synthetic scattering or a structured-light scan of a seafloor mock-up—and compare absolute pose error and depth error, with and without the 300-step test-pose refinement. If Swimm3R's advantage disappears under metric ground truth or shrinks when refinement is disabled, the central claim is not supported.","tokens_in":13732,"feed_emoji":"🌊","tokens_out":6606,"duration_ms":62449,"temperature":0.7,"pith_summary":"Swimm3R tries to establish that underwater 3D reconstruction fails not because geometry is absent, but because scattering and wavelength-dependent attenuation break the correspondence and color cues that standard structure-from-motion relies on. The paper's solution is to take a feed-forward SfM model pretrained on in-air images, fine-tune it on underwater images with a physics-based image-formation head, and use the resulting poses and restored dense point clouds to initialize a splatting renderer. The renderer uses Beta kernels and an explicit medium field so backscatter is modeled separately from scene surfaces. If the approach is right, turbid, low-texture underwater video can produce dense maps, accurate camera trajectories, and renderable scenes without per-scene calibration, and those maps directly improve visual localization. The claimed evidence includes better depth and pose accuracy on the FLSea test sequences and a 1.47 dB PSNR gain over WaterSplatting on a new Barbados dataset.","feed_headline":"Scattering-aware 3D pipeline beats prior underwater renderers by 1.47 dB","feed_subtitle":"Physics-informed SfM plus Beta splatting also raises pose accuracy on turbid Barbados video by 2–2.4 points.","key_machinery":"Two coupled pieces carry the argument. First, medium-aware SfM: LoRA fine-tuning of MASt3R-SfM with feature-map distillation plus an underwater physics head trained with forward, inverse, and saturation losses; the head regresses $\\beta_D$, $\\beta_B$, and $B_\\infty$ and implements $I = J\\odot T_D + B_\\infty\\odot(1-T_B)$ with $T_D=e^{-\\beta_D z}$, $T_B=e^{-\\beta_B z}$. Second, Underwater Beta Splatting: Beta primitives, an explicit medium field $(\\beta_B,\\beta_D,B_\\infty)(u)=\\sigma_{\\mathrm{med}}[A_\\phi(\\mathrm{SH}(d(u)))]$, and Scattering-aware Geometric Gradients, a gradient preconditioner that splits pixel-primitive adjoints into geometry and medium branches and rescales center steps for m","core_discovery":"The central claim is that underwater SfM can be made medium-aware by distilling in-air geometric priors into a feed-forward backbone and coupling it with a splatting renderer that explicitly models the water column. Swimm3R distills the encoder, decoder, and descriptor features of MASt3R-SfM into an underwater student via LoRA, and appends a physics head that predicts per-pixel direct attenuation $\\beta_D$, backscatter $\\beta_B$, and veiling light $B_\\infty$ from raw frames. The head implements the forward formation model $I = J \\odot e^{-\\beta_D z} + B_\\infty \\odot (1 - e^{-\\beta_B z})$ and its inverse restoration, so the student recovers clean radiance and geometry in one pass; at inferenc","pith_inferences":["The reported PSNR is measured after fitting each test camera pose to its held-out image, so it is a rendering-after-pose-fitting score, not pure novel-view synthesis; an evaluation without pose refinement would likely lower absolute numbers and could change the ranking.","The distillation recipe is not limited to water: any scattering or attenuating medium with a known forward model (fog, smoke, low light) could use the same LoRA-plus-physics-head transfer, provided paired clean-degraded data exist.","The GLUEMAP-anchored pose metrics on Barbados compare relative map quality; a controlled dataset with metric ground truth would be needed to know absolute accuracy.","The paper's water-type augmentation samples four veiling-light categories; testing on a fifth unseen water type would probe whether the physics head generalizes or merely memorizes the training palette."],"forward_implications":["Underwater robots could build dense, visually restored maps directly from raw video, without needing reliable sparse keypoints or prior calibration, which matters in turbid and low-texture environments.","Any downstream splatting renderer inherits a better initialization from Swimm3R's point clouds; the paper reports that replacing GLUEMAP or MASt3R-SfM initialization with Swimm3R improves rendering across renderers.","Explicitly separating backscatter and attenuation from surface radiance should make rendered maps stable under changing water conditions, supporting query-to-render matching for localization.","Because the SfM head also outputs optical parameters, the same framework can output scene structure and water-medium estimates simultaneously, useful for monitoring and navigation."],"supporting_citations":[{"why":"Base feed-forward SfM model whose in-air geometric priors are distilled into the underwater student via LoRA; provides the architecture and the FLSea evaluation baseline.","marker":"[19]"},{"why":"Supplies the LoRA feature-distillation strategy for adapting a feed-forward SfM model to a degraded-imaging domain, which Swimm3R transfers to underwater.","marker":"[32]"},{"why":"Introduces the Beta primitives that UWBS adopts as its scene representation, giving adaptive spatial and angular support.","marker":"[21]"},{"why":"The strongest underwater splatting baseline; its static medium MLP and rendering results are extended and compared against, with a reported 1.47 dB PSNR gain.","marker":"[14]"},{"why":"Provides the depth-dependent attenuation and backscatter image-formation model used for the physics head and augmentation.","marker":"[8]"},{"why":"Supplies the joint pose-primitive optimization and the test-time pose refinement protocol used for rendering and localization evaluation.","marker":"[36]"},{"why":"Base photometric loss and rasterization framework that underwater splatting methods, including UWBS, build on.","marker":"[11]"},{"why":"Dataset used for training and for depth/pose evaluation against COLMAP reference on clean sequences.","marker":"[39]"},{"why":"Feed-forward global SfM baseline and provider of reference poses for Barbados localization evaluation.","marker":"[37]"}],"fun_headline_variants":["Medium-aware SfM and Beta splatting lift underwater PSNR by 1.47 dB","Swimm3R: Underwater 3D improved 1.47 dB with Beta splatting","Beta splatting + medium-aware SfM: clear underwater 3D, +1.47 dB","Scattering-aware SfM boosts underwater localization by 2 points"],"cache_read_input_tokens":2816,"weakest_assumption_plain":"The load-bearing premise is that the Barbados evaluation numbers are trustworthy even though there is no metric ground truth for poses: the reference poses come from GLUEMAP, test poses are refined for 300 steps against the held-out image, and the reported gains could be optimistic if either the reference poses are wrong or the refinement protocol inflates render and localization scores.","fun_headline_variants_meta":{"raw":{"variants":["Medium-aware SfM and Beta splatting lift underwater PSNR by 1.47 dB","Swimm3R: Underwater 3D improved 1.47 dB with Beta splatting","Beta splatting + medium-aware SfM: clear underwater 3D, +1.47 dB","Scattering-aware SfM boosts underwater localization by 2 points"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001089,"raw_usage":{"total_tokens":4398,"prompt_tokens":769,"completion_tokens":3629,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":513,"completion_tokens_details":{"reasoning_tokens":3529}},"tokens_in":513,"tokens_out":3629,"duration_ms":24546,"temperature":1.0,"reasoning_tokens":3529,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T00:35:03.959558+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run Swimm3R, MASt3R-SfM, and WaterSplatting on an underwater scene with known metric poses and depth—for example a tank with calibrated cameras and synthetic scattering or a structured-light scan of a seafloor mock-up—and compare absolute pose error and depth error, with and without the 300-step test-pose refinement. If Swimm3R's advantage disappears under metric ground truth or shrinks when refinement is disabled, the central claim is not supported.","supporting_citations":[{"cited_title":"MASt3R-SfM: a fully-integrated solution for unconstrained structure-from-motion,","cited_arxiv_id":null,"evidence_quote":"Base feed-forward SfM model whose in-air geometric priors are distilled into the underwater student via LoRA; provides the architecture and the FLSea evaluation baseline."},{"cited_title":"Universal Beta Splatting,","cited_arxiv_id":null,"evidence_quote":"Introduces the Beta primitives that UWBS adopts as its scene representation, giving adaptive spatial and angular support."},{"cited_title":"3D Gaussian Splatting for real-time radiance field rendering","cited_arxiv_id":null,"evidence_quote":"Base photometric loss and rasterization framework that underwater splatting methods, including UWBS, build on."},{"cited_title":"Global Structure-from-Motion meets feedforward reconstruction,","cited_arxiv_id":null,"evidence_quote":"Feed-forward global SfM baseline and provider of reference poses for Barbados localization evaluation."}],"review_version":1}