{"id":"3b53695d-e6f2-43a7-91ec-4a6b72783567","arxiv_id":"2411.13230","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":5,"one_line_summary":"OceanLens, a neural network combining physics-based backscatter and attenuation estimation with Sobel and Laplacian-of-Gaussian losses, reports large improvements in GPMAE and UIQM, though on a small set of images with unsupported headline numbers.","lead":"OceanLens is a deep learning system that enhances underwater photographs by separately removing scattered light and correcting color loss from depth-dependent light absorption. The authors report large metric gains, but the evidence is thin, with only a handful of test images and headline numbers that do not always match the tables.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Aggregate performance claims in the abstract (65% GPMAE, 60% UIQM, 12-15% SSIM) are not derivable from Tables I-V, so the empirical core of the paper is unsupported.","rationale":"I reviewed the strongest claim as an empirical performance advantage, since physics-based modeling is not novel in itself. For this claim to hold, the paper must provide a reproducible accounting of the reported averages. That accounting is missing: Table V contains three images and no 12-15% SSIM result; Tables I-IV combine single values for D1-D4 with multiple per-patch values for D5 and include known failure cases; the conclusion repeats the abstract numbers without derivation. This is not a dispute with consensus or a novelty issue; it is an internal evidentiary gap. The reader's depth-map concern is real but secondary: validating MonoDepth2 and Depth-Anything-V2 depths against SeeThru SfM depth would test generalizability, but it would not repair the unsubstantiated aggregate metrics. Thus I maintain the reader's REJECT verdict; no change is needed. A single analytical recomputation of the headline percentages from the tables would settle the issue.","tokens_in":11579,"tokens_out":5503,"duration_ms":56048,"concrete_test":"Independently recompute the abstract's three headline figures from the published tables: define the per-image GPMAE reduction relative to ST and DSC (or to raw, if that is the intended baseline) for Tables I and III, average over D1-D5, and do the same for UIQM in Tables II and IV and SSIM in Table V. If no explicit, reproducible aggregation rule yields 65%, 60%, and 12-15%, the central empirical claim is unsupported; if the authors' released code and checkpoints reproduce the tables under the stated rule, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is an empirical improvement over SeeThru and DeepSeeColor. The paper's own tables do not substantiate the advertised aggregates. Table V reports only three UIEB images; relative SSIM changes are +865% (Image1), a sign-crossing increase from -0.0062 to 0.0308 (Image2), and +6.3% (Image3). No entry is 12-15%. For SeeThru, D1-D4 show large GPMAE reductions, but D5 is a documented failure with values exceeding ST/DSC, and Tables I and III contain repeated per-patch entries for D5 only, so it is unclear what aggregation produced the 65% average. UIQM gains are also inconsistent: DA on D1 and D5 decreases UIQM relative to raw (Tables II and IV), and no stated averaging rule reproduces exactly 60%. The abstract's 65%/60%/12-15% figures are the load-bearing content; since the provided tables and text do not permit recomputing them, the claimed advantage is unverified. The depth-map assumption identified by the reader is secondary: even if MonoDepth2/DA depths are accurate, the headline numbers would still not follow from the evidence.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript proposes OceanLens, a two-branch neural network for underwater image enhancement that estimates backscatter and deattenuation from an image and a depth map, using an adaptive Huber backscatter loss and a composite deattenuation loss with Sobel and LoG edge terms. The method is tested on SeeThru (5 images), US Virgin Islands (qualitative only), and UIEB (3 images), with claims of 65% GPMAE reduction, 60% UIQM gain, and 12-15% SSIM improvement over SeeThru and DeepSeeColor.","tokens_in":11829,"tokens_out":3549,"duration_ms":33885,"significance":"If the claims held, the work would provide a useful lightweight enhancement method and evidence that terrestrial monocular depth estimators transfer to underwater scenes. The authors also make code available. However, the reported aggregate numbers are not supported by the tables, the main quality metric is partially aligned with the training objective, and the evaluation set is too small to justify the claims.","major_comments":[{"comment":"The abstract and conclusion state a 65% GPMAE reduction, 60% UIQM increase, and 12-15% SSIM improvement, but these aggregates are not derivable from Tables I-V. Table V reports only three UIEB images, with relative SSIM changes of +865% (Image1), a sign-crossing increase (Image2 from -0.0062 to 0.0308), and +6.3% (Image3); no entry is 12-15%. Tables I and III contain repeated per-patch entries for D5 only, and D5 shows GPMAE values exceeding ST/DSC and UIQM values below raw, yet the averaging rule for the 65% and 60% figures is not stated. The authors need to report per-image results, define the aggregation, and include variability estimates across training runs.","section":"Abstract and Section IV-B"},{"comment":"There is a circularity concern: Lsat with Isat_tar=1 and Lint with Itar push the output toward white/gray, while the GPMAE metric of Eq. (21) measures angular error against the (1,1,1) white vector. Minimizing these losses therefore directly reduces the reported evaluation metric by construction, so GPMAE improvements do not independently demonstrate color fidelity. The authors should either use a reference-based metric (e.g., on UIEB) for the SeeThru comparison or train without the saturation/intensity targets and show GPMAE still improves.","section":"III-C.1, Eqs. (14)-(15) and IV-A.1, Eq. (21)"},{"comment":"The claim that pre-trained terrestrial monocular depth models (MonoDepth2, Depth-Anything-V2-Large) are suitable for underwater backscatter estimation is not validated. No quantitative comparison of MD/DA depth maps against the SeeThru ground-truth depth is provided, and Tables I-IV show large performance variations across depth sources (e.g., DA GPMAE 0.57 on D3 but 20-28 on D5 in Table III). A depth accuracy evaluation and its correlation with enhancement performance are needed before this assumption can support the method.","section":"Section I and IV-B"},{"comment":"The evaluation uses only five SeeThru images and three UIEB images, with no error bars, multiple runs, or statistical tests. The D5 failure is acknowledged in Section IV-B but is excluded from the aggregate claims without a stated rule. This sample size is insufficient to support the general superiority claims in the abstract; per-image results with variance and a pre-registered aggregation rule are required.","section":"Tables I-V"}],"minor_comments":[{"comment":"The word 'deattnuation' in the section heading should be 'deattenuation'.","section":"Section III-C"},{"comment":"The heading 'Evalution Metrics' contains a typo; it should be 'Evaluation Metrics'.","section":"Section IV-A"},{"comment":"The D5 rows contain slash-separated values (e.g., 17/16/15/17) without explanation; clarify whether these are repeated patch measurements, different runs, or something else.","section":"Tables I and III"},{"comment":"The caption contains 'OceanLeans', which should be 'OceanLens'.","section":"Fig. 7 caption"},{"comment":"The UIEB results are reported for only three images with no image identifiers; provide the image names or indices so that the results are reproducible.","section":"Section IV-B"},{"comment":"Training hyperparameters (learning rate, optimizer, number of epochs, P, T, and the chosen values of delta and beta) are not specified; include them to allow replication.","section":"Sections III-B and III-C"}],"recommendation":"reject","confidential_remarks":"The discrepancy between the abstract's headline numbers and the tables is severe, and the GPMAE metric is partly optimized by the loss design, which undermines the main evaluation. I would not consider this paper publishable in its current form; a major rework with a larger evaluation and independent metrics would be needed."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is an honest, clearly written incremental extension of DeepSeeColor with two added edge-loss terms and an adaptive Huber backscatter term. The paper explicitly credits DeepSeeColor and reuses its EAF/CEAF activations plus three of five loss components. That is legitimate. The Sobel/LoG plus adaptive Huber combination is new as a package, though each ingredient is standard. The authors also test three depth-map sources, including two off-the-shelf monocular estimators, and they acknowledge a failure case (D5). That is more transparent than much of the literature.\n\nWhat the paper does not do is support its abstract. The 65% GPMAE, 60% UIQM, and 12-15% SSIM figures are not derivable from the tables. Table V shows three UIEB images; relative SSIM changes are +865% (from 0.0428 to 0.413, but on a tiny absolute base), a sign-crossing move from -0.0062 to 0.0308, and +6.3%. No entry is 12-15%, and the text itself notes Image2 and Image3 PSNR drops. On SeeThru, D5 is a documented failure with multiple per-patch entries scattered across Tables I/III, so the averaging rule that produces a clean 65% is unclear. UIQM scores in Tables II/IV show several decreases (e.g., DA on D1 and D5) relative to raw, so the 60% increase is not reproducible from the reported numbers.\n\nThere is also a metric-design circularity worth flagging: the saturation and intensity targets in the loss push output pixels toward white/gray, and GPMAE measures angular distance to (1,1,1). Part of the reported GPMAE improvement is the model being trained against the metric. That does not make the loss useless, but it means the headline metric is partly self-fulfilling.\n\nThe evaluation is tiny: five SeeThru images, three UIEB images, no error bars, no training details (epochs, learning rate, split), and hyperparameters (delta, layer count) chosen on the same images used for scoring. The depth-map claim is plausible but unvalidated; no comparison against ground-truth depth is provided, and results vary widely across depth sources.\n\nWho this is for: someone working on underwater image enhancement who wants a list of loss tweaks to try, or a teaching example of why aggregate claims need to be recomputable. It is not a paper whose empirical conclusions can be trusted as stated.\n\nRecommendation: send to peer review for major revision, not desk reject. The method is clearly stated, the failure case is acknowledged, and there is a code link. A serious referee can force the authors to either produce the actual averages, add error bars, and evaluate on a larger subset, or drop the aggregate claims. If they cannot, reject. The paper deserves that chance.","headline":"A modest loss-function extension of DeepSeeColor whose headline numbers do not survive contact with its own tables.","tokens_in":12405,"tokens_out":4055,"would_cite":false,"duration_ms":37820,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"OceanLens claims that splitting an underwater image into estimated backscatter and attenuation, with Sobel and LoG edge losses, cuts GPMAE by 65% and raises UIQM by 60% over SeeThru and DeepSeeColor.","keywords":["underwater image enhancement","backscatter estimation","monocular depth estimation","edge-preserving loss","Sobel filter","Laplacian of Gaussian","UIQM","GPMAE"],"falsifier":"Take a SeeThru image and run OceanLens with its original SfM depth map, with MonoDepth2's predicted map, with Depth-Anything-V2-Large's predicted map, and with a deliberately scrambled depth map; if GPMAE and UIQM change little across all four conditions, the claim that depth drives the enhancement fails, whereas large degradation under scrambling would support it.","tokens_in":11349,"feed_emoji":"🌊","tokens_out":7893,"duration_ms":75457,"temperature":0.7,"pith_summary":"OceanLens is a deep-learning method for restoring color and detail in underwater photographs. The paper argues that by splitting the image into two physically meaningful parts—backscattered light and depth-dependent attenuation—and estimating each with its own convolutional network, enhancement can beat two published approaches (SeeThru and DeepSeeColor) without requiring multi-view reconstruction. The network is trained with an adaptive Huber backscatter loss and two edge-preserving losses (Sobel and Laplacian of Gaussian) that protect fine structure. The paper reports an average 65% reduction in Grayscale Patch Mean Angular Error and a 60% increase in the Underwater Image Quality Metric, and shows that adding convolution layers improves SSIM on the UIEB reference dataset. The broader promise is that a single image plus an off-the-shelf monocular depth estimate is enough to drive the correction.","feed_headline":"Neural net cuts underwater color error by 65 percent","feed_subtitle":"The model splits images into backscatter and attenuation, then uses Sobel and LoG edge losses to keep detail.","key_machinery":"The load-bearing mechanism is the split of the underwater image into backscatter and attenuated direct signal, following the formation model $I^c = I_D^c + I_B^c$. A backscatter network turns the range/depth map into $\\hat{I}_B$ via convolutional terms with CEAF/EAF activations; a deattenuation network turns the same map into $\\hat{\\alpha}_D^c(z)$, a sum of exponentials. The enhanced image is the product of the direct estimate and inverse attenuation. The objective that stabilizes this decomposition is a composite loss: an adaptive Huber backscatter term, saturation/intensity/variance terms, and Sobel plus Laplacian-of-Gaussian edge terms, which bias the network toward preserving luminance variance and high-frequency detail.","core_discovery":"On its own terms, the paper's discovery is that a physics-grounded network pair—one estimating backscatter $\\hat{I}_B$ from a range map, the other estimating inverse attenuation $\\hat{\\alpha}_D^c(z)$—can reconstruct a direct image $\\hat{I}_J = (I - \\hat{I}_B)\\,\\hat{\\alpha}_D^c(z)$ with better color and structure than the SeeThru and DeepSeeColor baselines. The backscatter estimate follows the exponential form $I_B = I_{B\\infty}(1-\\exp(-b_1 z)) + I'_B \\exp(-b_2 z)$, implemented with complementary exponential and exponential activations; the attenuation estimate is a neural sum of exponentials. Three additions carry the reported gains: the adaptive Huber loss on backscatter, the Sobel and LoG edge losses in the deattenuation objective, and extra convolutional layers for detail. The paper also claims that depth maps from terrestrial-trained monocular models—MonoDepth2 and Depth-Anything-V2-Large—deliver performance in line with the original structure-from-motion range maps, with the larger Depth-Anything-V2-Large model giving the lowest angular errors on several SeeThru images.","pith_inferences":["If terrestrial-trained monocular depth estimates truly suffice, then any single underwater image without range data becomes a candidate for physics-based enhancement; this could be tested by comparing OceanLens output using predicted depth against output using the SeeThru ground-truth range map on the same images.","The depth-source sensitivity visible in the reported tables suggests a natural next benchmark: matching the depth estimator to water conditions and lighting, since Depth-Anything-V2-Large wins on some images while the original depth map wins on others.","The backscatter-plus-attenuation split is not unique to water, so the same network structure with edge losses could be transferred to fog, haze, or turbid-media restoration by retraining the coefficient ranges.","The Sobel and LoG edge-loss combination could serve as a generic structural regularizer for image restoration, because it penalizes gradient mismatch and second-derivative mismatch separately."],"forward_implications":["OceanLens reports an average 65% lower Grayscale Patch Mean Angular Error and a 60% higher Underwater Image Quality Metric than the SeeThru and DeepSeeColor baselines on the SeeThru images.","Depth maps from MonoDepth2 and Depth-Anything-V2-Large can replace structure-from-motion range maps for the backscatter and attenuation estimates, with Depth-Anything-V2-Large giving the lowest angular errors on several images.","Adding convolutional layers improves structural fidelity: SSIM on the UIEB reference images rises, with Image3 up 6.3% and Image2 moving from negative to 0.0308, described as up to 12-15% overall SSIM improvement.","The Sobel and LoG edge losses measurably improve UIQM in the ablations, and the method runs fast enough for near real-time use at roughly 4-5 milliseconds per 7-12 MB image."],"supporting_citations":[{"why":"Supplies the revised underwater image formation model, the SeeThru dataset with original depth maps, and the physical equations the networks implement.","marker":"[4]"},{"why":"Defines the SeeThru baseline method that OceanLens is compared against and aims to replace.","marker":"[7]"},{"why":"DeepSeeColor is the comparison baseline and the source of the CEAF/EAF activations and network structure OceanLens extends.","marker":"[9]"},{"why":"MonoDepth2 provides one of the two terrestrial-trained monocular depth maps tested as a substitute for structure-from-motion range maps.","marker":"[10]"},{"why":"Depth-Anything-V2-Large provides the second monocular depth source and yields the lowest GPMAE values on several SeeThru images.","marker":"[11]"},{"why":"The UIEB benchmark supplies reference images against which PSNR and SSIM are computed for the multi-layer version.","marker":"[26]"},{"why":"Defines the Gray Patch Mean Angular Error metric used to report the 65% color-error reduction.","marker":"[12]"},{"why":"Defines the Underwater Image Quality Measure used to report the 60% quality improvement.","marker":"[39]"},{"why":"Supplies the adaptive Huber regression loss used as the backscatter loss function.","marker":"[36]"}],"fun_headline_variants":["Underwater imaging AI cuts color error by 65%","AI model improves underwater color, cuts error 65%","OceanLens deep learning reduces underwater error by 65%","Underwater photos: AI restores color, drops error 65%"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The method's results depend on pre-trained monocular depth models, trained on ordinary land images, producing depth maps that are accurate enough for underwater backscatter estimation, and the paper does not validate those predicted depths against ground-truth underwater range maps.","fun_headline_variants_meta":{"raw":{"variants":["Underwater imaging AI cuts color error by 65%","AI model improves underwater color, cuts error 65%","OceanLens deep learning reduces underwater error by 65%","Underwater photos: AI restores color, drops error 65%"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001169,"raw_usage":{"total_tokens":4889,"prompt_tokens":1053,"completion_tokens":3836,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":669,"completion_tokens_details":{"reasoning_tokens":3765}},"tokens_in":669,"tokens_out":3836,"duration_ms":27390,"temperature":1.0,"reasoning_tokens":3765,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T16:41:04.060345+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a SeeThru image and run OceanLens with its original SfM depth map, with MonoDepth2's predicted map, with Depth-Anything-V2-Large's predicted map, and with a deliberately scrambled depth map; if GPMAE and UIQM change little across all four conditions, the claim that depth drives the enhancement fails, whereas large degradation under scrambling would support it.","supporting_citations":[{"cited_title":"A revised underwater image formation model,","cited_arxiv_id":null,"evidence_quote":"Supplies the revised underwater image formation model, the SeeThru dataset with original depth maps, and the physical equations the networks implement."},{"cited_title":"Sea-thru: A method for removing water from underwater images,","cited_arxiv_id":null,"evidence_quote":"Defines the SeeThru baseline method that OceanLens is compared against and aims to replace."},{"cited_title":"Deepseecolor: Realtime adaptive color correction for autonomous underwater vehicles via deep learning methods,","cited_arxiv_id":null,"evidence_quote":"DeepSeeColor is the comparison baseline and the source of the CEAF/EAF activations and network structure OceanLens extends."},{"cited_title":"An underwater image enhancement benchmark dataset and beyond,","cited_arxiv_id":null,"evidence_quote":"The UIEB benchmark supplies reference images against which PSNR and SSIM are computed for the multi-layer version."},{"cited_title":"Diving into haze-lines: Color restoration of underwater images,","cited_arxiv_id":null,"evidence_quote":"Defines the Gray Patch Mean Angular Error metric used to report the 65% color-error reduction."},{"cited_title":"Adaptive huber regression,","cited_arxiv_id":null,"evidence_quote":"Supplies the adaptive Huber regression loss used as the backscatter loss function."}],"review_version":1}