REVIEW 4 major objections 6 minor 1 cited by
Marine Snow Removal Using Internally Generated Pseudo Ground Truth
T0 review · 4 major / 6 minor · reviewed 2026-08-16 · deepseek-v4-flash
Pith's one-line read Real marine snow masks extracted from video can train an enhancer that improves underwater SLAM.
desk verdict A genuinely useful pseudo-pair trick for marine snow training data, with a SLAM evaluation that is directionally right but statistically too thin to fully land the headline claim. read the letter →
The pith
A machine-rendered reading of the paper's core claim, the machinery that carries it, and where it could break.
The reading
What carries the argument
The mechanism is median-subtracted snow-mask extraction and overlay. A median frame over $N$ frames estimates the static background in a snow-affected clip; subtracting it from each frame leaves only the dynamic bright particles, which are then overlaid onto clean reference patches at randomised temporal and spatial offsets to form paired training data. The second load-bearing component is BVI-Mamba, a video enhancement architecture with feature-level frame alignment and 2D Selective Scan state-space modelling, which is trained with an $\ell^1$ loss to map snow-affected frames to the clean references.
What would settle it
Take a snow-free static scene, introduce a small continuous camera pan during the clip, apply the paper's median-subtraction pipeline, and train the enhancement network on the resulting pairs; if the network systematically removes or blurs stationary edges rather than only particles, the mask is contaminated. A second check is to measure mask purity directly by recording the same underwater scene with two exposures, using the short exposure to identify true snow particles and computing what fraction of the extracted mask pixels overlap true particles.
Extended reading notes
Core claim
The central claim is that real marine snow artefacts extracted by median-frame subtraction can serve as high-quality training signal, and that a network trained on these internally generated pairs generalises to unseen footage and improves feature-based underwater SLAM. The assertion is that because the training masks come from real video rather than synthetic simulation, the model learns the true appearance and motion of marine snow, and because enhancement acts on entire frames rather than by rejecting keypoints, the downstream mapping pipeline needs no modification. The evidence is a comparative evaluation on eight real sequences where the proposed enhancement yields a lower number of detected keypoints per frame, interpreted as fewer snow detections, a consistently higher number of frame-to-frame feature matches, and a denser 3D reconstruction, with the average dense-map point count rising about 28% over the no-enhancement baseline.
Load-bearing premise
The load-bearing premise is that the median of a frame window isolates marine snow, which holds only if the scene behind the snow is static for the whole window and the only moving signal is snow; any camera drift, vehicle motion, or passing non-snow object contaminates the mask with real scene structure that the network is then trained to erase.
Editorial extensions
If this is right
- Preprocessing underwater inspection footage with this enhancement should increase the number of valid frame-to-frame feature matches, improving feature-tracking stability in SLAM.
- Dense 3D reconstruction should become more complete, since the paper's average reconstructed point count rises from 893,076 to 1,142,521.5 with enhancement.
- Because enhancement is applied to entire frames rather than by filtering keypoints, the same preprocessing can benefit any downstream underwater vision task that consumes raw imagery, such as object detection or inspection.
- The dataset-generation procedure is portable: it only needs a snow-affected clip and a clean reference patch, so it can be re-run on footage from different sites, depths, or cameras to build customised paired datasets.
Reading between the lines
- If the median-subtraction masks are as clean as assumed, the method turns unlabelled underwater video into effectively unlimited paired training data; a natural extension is to apply the same pipeline to other transient visual noise such as bubbles, sediment plumes, or drifting algae that move against a static background.
- Part of the dense-map point increase could come from the network smoothing low-texture seabed and generating spurious matches rather than from recovering true structure; a test on a synthetic scene with known geometry and simulated snow would separate snow suppression from detail hallucination.
- The approach is likely sensitive to camera motion and vehicle manoeuvring, since the static-background assumption degrades as motion grows; a controlled sweep of camera translation speed would quantify the ceiling of the method.
Signed reviews
Editorial analysis
A structured set of objections, weighed in public.
Referee Report
Summary. The paper proposes a self-supervised data-generation pipeline for marine snow removal. Static background patches are selected from raw underwater videos, a temporal median frame is subtracted from each frame to obtain dynamic snow masks (Eq. 2), and these masks are additively overlaid onto clean reference frames to create paired snowy/snow-free training data (Eq. 3). The authors then train an enhancement network based on BVI-Mamba on these internally generated pseudo-pairs and evaluate the enhanced videos by feeding them into a proprietary SLAM and dense-mapping system. The reported claims are that the method reduces snow-like keypoints, consistently increases frame-to-frame feature matches, and raises the average dense-map point count from 893,076 to 1,142,521.5 across eight real sequences.
Significance. If the central claim is substantiated, the paper makes a useful practical contribution: it avoids fully synthetic snow simulation by reusing real snow artifacts, enables supervised training without manually cleaned ground truth, and evaluates the enhancement through downstream SLAM metrics rather than pixel-level metrics that the authors rightly argue are not meaningful for this task. The idea of internally generated pseudo-pairs from median subtraction is simple and potentially transferable to other transient-noise removal problems. The main value is therefore conditional on whether the reported SLAM improvement is statistically real and whether the mask-extraction assumptions hold in the deployment conditions. The paper does not provide machine-checked proofs or code, but the described pipeline is straightforward enough that the missing per-sequence statistics and mask-purity measurements could be supplied in a revision.
major comments (4)
- [Section IV-C, Table I] The headline quantitative claim rests entirely on a single average dense-map point count over eight sequences, with no standard deviation, no per-sequence values, and no paired significance test. With n = 8, one favorable sequence could move the average by roughly the reported 28% margin, so the claimed improvement from 893,076 to 1,142,521.5 points is not yet supported. Please report the per-sequence point counts, the mean and spread for each method, and a paired statistical test (e.g., Wilcoxon signed-rank or a bootstrap confidence interval on the mean difference). Without this, the central claim that enhancement improves SLAM is not independently auditable.
- [Section III-A, Eq. (2)] The mask-extraction step assumes that the scene is static over the median window and that marine snow is the only dynamic signal. If the camera drifts, the vehicle moves, or bubbles, fish, or other objects pass through the selected patch, the subtraction in Eq. (2) leaves scene residuals in the mask, and the network is trained to erase real structure. Manual patch selection reduces this risk but does not quantify it. The paper should provide a quantitative purity analysis of the extracted masks, for example residual energy outside detected snow components, or a comparison of masks with manually annotated snow regions, and should state how sensitive the results are to the median window length N and patch size.
- [Section IV-C, Figs. 4 and 6] The claim that the method 'consistently achieves a higher number of frame-to-frame matches' is supported only by smoothed curves without aggregate statistics, per-sequence breakdowns, or error bars. Figure 6 in particular cannot substantiate 'consistently' when the underlying per-sequence values are not shown. Please provide a table of frame-to-frame match counts (or match-rate ratios) for each of the eight sequences, with means and confidence intervals, and state how the smoothing was applied and over what window. The same applies to the keypoint-count claim in Fig. 4, where 'lower is better' is asserted but no statistical summary is given.
- [Section IV-B and Section IV-C] The paper correctly argues that pixel-level metrics are not the right evaluation for this task, but it then substitutes a single proprietary SLAM/dense-mapping pipeline and a small set of eight sequences as the only evidence. Because the SLAM system and the underwater sequences are proprietary, the result cannot be independently reproduced. At minimum, the authors should report the variance across sequences and ideally release the per-sequence metrics or a runnable evaluation protocol. This is load-bearing because the entire contribution is the claimed improvement in downstream task performance.
minor comments (6)
- [Abstract] The phrase 'snow, free underwater videos' should be 'snow-free underwater videos', and 'with the absence of groundtruth' should be rephrased as 'in the absence of ground truth'.
- [Section I] The phrase 'loosing details' should be 'losing details'.
- [Section III-A, Eq. (3)] The index notation is confusing: I_t^snow is indexed by t, while the clean frame is indexed by t' and the spatial offsets are Delta x, Delta y. Please define the ranges of t, t', and the sliding window, and clarify whether the clean frame and the mask are temporally aligned within the window or chosen independently.
- [Section III-B and References] The architecture is called BVI-Mamba but the cited reference [18] is titled 'BVI-RLV: A fully registered dataset and benchmarks for low-light video enhancement'. If BVI-Mamba is a distinct method or is detailed elsewhere, please cite the correct source or clarify the relationship.
- [Figure 7, Section IV-C] The text describes two reconstructed sequences (2023 000015 000135 and 2022 005644 005744) but the figure caption names only 2022 004742 004844. Please align the caption with the sequences actually shown and specify which point clouds correspond to which sequence.
- [Section IV-A] The patch selection for static background regions is described only as 'fixed-size patches of 550 x 600 pixels' that 'avoid any foreground objects'. Please state how many patches were selected, from how many source sequences, and what manual criteria were used, so that the dataset composition is reproducible.
Circularity Check
No significant circularity; the internally-generated pseudo-GT is transparent and the headline SLAM claim is grounded in external evaluation on real sequences.
full rationale
The paper's pipeline is explicitly a synthetic-pair construction: Eq. (1)-(2) extract a dynamic mask by temporal median subtraction, and Eq. (3) overlays that mask on manually selected low-snow frames to form training pairs. The enhancement network is trained to invert Eq. (3), but this is the stated dataset-generation design, not a hidden derivation. The paper does not use pixel-level restoration as its evidence; Section IV-B explicitly disclaims PSNR/SSIM and admits the 'ground truth' is only the cleanest available data, not truly clean. The central claim—improved SLAM feature matching and denser 3D reconstruction—is tested on eight real underwater sequences with Beam's proprietary mapping system (Section IV-C), which is external to the training objective and to the pseudo-GT construction. The cited prior works (BVI-Mamba [18], BEM [15]) serve as architecture and baseline, not as load-bearing proofs. The main weakness is statistical—single averaged point counts in Table I, smoothed curves in Figs. 4 and 6, no per-sequence breakdown or variance—which is an evaluation-rigor concern, not circularity. Therefore no specific circular step is identified.
Assumptions & free parameters
free parameters (2)
- Median window length N =
not reported
- Static background patch size =
550x600 pixels
assumptions (4)
- domain assumption Temporal median subtraction isolates moving snow from a static background.
- domain assumption Additive overlay of snow masks onto clean frames reproduces realistic snow corruption.
- domain assumption A network trained on pseudo-pairs generalizes to real marine snow.
- domain assumption SLAM metrics, including dense point count, reflect enhancement quality.
Cite this review
Pith. "Pith review of Marine Snow Removal Using Internally Generated Pseudo Ground Truth." pith.science (2026). https://pith.science/paper/7FYPVQDP
@misc{pith2026250419289,
author = {Pith},
title = {Pith review of: Marine Snow Removal Using Internally Generated Pseudo Ground Truth},
year = {2026},
howpublished = {\url{https://pith.science/paper/7FYPVQDP}},
note = {Machine review of arXiv:2504.19289}
}
read the original abstract
Underwater videos often suffer from degraded quality due to light absorption, scattering, and various noise sources. Among these, marine snow, which is suspended organic particles appearing as bright spots or noise, significantly impacts machine vision tasks, particularly those involving feature matching. Existing methods for removing marine snow are ineffective due to the lack of paired training data. To address this challenge, this paper proposes a novel enhancement framework that introduces a new approach for generating paired datasets from raw underwater videos. The resulting dataset consists of paired images of generated snowy and snow, free underwater videos, enabling supervised training for video enhancement. We describe the dataset creation process, highlight its key characteristics, and demonstrate its effectiveness in enhancing underwater image restoration in the absence of ground truth.
Figures
Figures from the paper (4 more)
Forward citations
Cited by 1 Pith paper
-
Bayesian Neural Networks for One-to-Many Mapping in Image Enhancement
A two-stage BNN-DNN framework models one-to-many image enhancement by sampling diverse coarse illuminations and refining them with a deterministic network.
Reference graph
Works this paper leans on
-
[10]
Marine snow simulation and elimination in video,
J. P. Coffelt, N. Nowald, and P. Kampmann, “Marine snow simulation and elimination in video,” in 2023 IEEE Underwater Technology (UT) . IEEE, 2023, pp. 1–10
work page 2023
-
[17]
Detecting and suppressing marine snow for underwater visual slam,
L. M. Hodne, E. Leikvoll, M. Yip, A. L. Teigen, A. Stahl, and R. Mester, “Detecting and suppressing marine snow for underwater visual slam,” in Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, 2022, pp. 5101–5109
work page 2022
-
[1]
An improved guidance image based method to remove rain and snow in a single image,
J. Xu, W. Zhao, P. Liu, and X. Tang, “An improved guidance image based method to remove rain and snow in a single image,” Computer and Information Science , vol. 5, no. 3, p. 49, 2012
work page 2012
-
[2]
Content-adaptive rain and snow removal algorithms for single image,
S. Yu, Y . Zhao, Y . Mou, J. Wu, L. Han, X. Yang, and B. Zhao, “Content-adaptive rain and snow removal algorithms for single image,” in Advances in Neural Networks–ISNN 2014: 11th International Sympo- sium on Neural Networks, ISNN 2014, Hong Kong and Macao, China, November 28-December 1, 2014. Proceedings 11 . Springer, 2014, pp. 439–448
work page 2014
-
[3]
Removing rain and snow in a single image using saturation and visibility features,
S.-C. Pei, Y .-T. Tsai, and C.-Y . Lee, “Removing rain and snow in a single image using saturation and visibility features,” in 2014 IEEE International Conference on Multimedia and Expo Workshops (ICMEW). IEEE, 2014, pp. 1–6
work page 2014
-
[4]
Rain or snow detection in image sequences through use of a histogram of orientation of streaks,
J. Bossu, N. Hautiere, and J.-P. Tarel, “Rain or snow detection in image sequences through use of a histogram of orientation of streaks,” International journal of computer vision , vol. 93, pp. 348–367, 2011
work page 2011
-
[5]
Single image rain and snow removal via guided l0 smoothing filter,
X. Ding, L. Chen, X. Zheng, Y . Huang, and D. Zeng, “Single image rain and snow removal via guided l0 smoothing filter,” Multimedia Tools and Applications , vol. 75, pp. 2697–2712, 2016
work page 2016
-
[6]
Removing rain and snow in a single image using guided filter,
J. Xu, W. Zhao, P. Liu, and X. Tang, “Removing rain and snow in a single image using guided filter,” in 2012 IEEE International Conference on Computer Science and Automation Engineering (CSAE) , vol. 2. IEEE, 2012, pp. 304–307
work page 2012
Show all 19 references
-
[7]
Video snow removal based on self-adaptation snow detection and patch-based gaussian mixture model,
B. Yang, Z. Jia, J. Yang, and N. K. Kasabov, “Video snow removal based on self-adaptation snow detection and patch-based gaussian mixture model,” IEEE Access , vol. 8, pp. 160 188–160 201, 2020
2020
-
[8]
Online rain/snow removal from surveillance videos,
M. Li, X. Cao, Q. Zhao, L. Zhang, and D. Meng, “Online rain/snow removal from surveillance videos,” IEEE Transactions on Image Pro- cessing, vol. 30, pp. 2029–2044, 2021
2021
-
[9]
Snow re- moval in video: A new dataset and a novel method,
H. Chen, J. Ren, J. Gu, H. Wu, X. Lu, H. Cai, and L. Zhu, “Snow re- moval in video: A new dataset and a novel method,” in 2023 IEEE/CVF International Conference on Computer Vision (ICCV) . IEEE, 2023, pp. 13 165–13 176
2023
-
[11]
Underwater image enhancement and marine snow removal for fishery based on integrated dual-channel neural network,
Y . Wang, X. Yu, D. An, and Y . Wei, “Underwater image enhancement and marine snow removal for fishery based on integrated dual-channel neural network,” Computers and Electronics in Agriculture , vol. 186, p. 106182, 2021. [Online]. Available: https://www.sciencedirect.com/sci...
2021
-
[12]
An underwater image enhancement benchmark dataset and beyond,
C. Li, C. Guo, W. Ren, R. Cong, J. Hou, S. Kwong, and D. Tao, “An underwater image enhancement benchmark dataset and beyond,” IEEE Transactions on Image Processing , vol. 29, pp. 4376–4389, 2020
2020
-
[13]
A deep learning approach for marine snow synthesis and removal,
F. Galetto and G. Deng, “A deep learning approach for marine snow synthesis and removal,” Signal, Image and Video Processing , vol. 19, no. 1, p. 1, 2025
2025
-
[14]
Underwater image restoration based on marine snow removal techniques,
Y . Li, B. Niu, T. Zhou, and H. Gao, “Underwater image restoration based on marine snow removal techniques,” in 2024 2nd International Conference on Signal Processing and Intelligent Computing (SPIC) , 2024, pp. 651–655
2024
-
[15]
Bayesian neural networks for one-to-many mapping in image enhancement,
G. Huang, N. Anantrasirichai, F. Ye, Z. Qi, R. Lin, Q. Yang, and D. Bull, “Bayesian neural networks for one-to-many mapping in image enhancement,” arXiv preprint arXiv:2501.14265 , 2025
2025 arXiv
-
[16]
Real-time marine snow noise removal from underwater video sequences,
B. Cyganek and K. Gongola, “Real-time marine snow noise removal from underwater video sequences,” Journal of Electronic Imaging , vol. 27, no. 4, pp. 043 002–043 002, 2018
2018
-
[18]
Bvi-rlv: A fully registered dataset and benchmarks for low-light video enhancement,
R. Lin, N. Anantrasirichai, G. Huang, J. Lin, Q. Sun, A. Malyugina, and D. R. Bull, “Bvi-rlv: A fully registered dataset and benchmarks for low-light video enhancement,” arXiv preprint arXiv:2407.03535 , 2024
2024 arXiv
-
[19]
Mamba: Linear-time sequence modeling with selective state spaces,
A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,” arXiv preprint arXiv:2312.00752 , 2023
2023 arXiv
Reviewed August 16, 2026 · model on record in the stance chip above.
Discussion (0). Continue with ORCID to comment.