{"id":"fca19ea0-4808-47d1-bd6e-1d258739dcef","arxiv_id":"2411.18296","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"HUPE is an invertible-network underwater enhancement method that combines frequency-domain affine coupling, dark-channel priors, and semantic feature collaboration to improve both image quality and downstream detection and segmentation.","lead":"HUPE is a deep learning system that uses a reversible neural network, physical depth and gradient cues, and semantic feature collaboration to enhance underwater images while training for object detection and segmentation. It claims better visual quality and better downstream perception than eight existing underwater enhancement methods.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 2's EUVP UIQM value for HUPE (20.9151) is an order-of-magnitude outlier that makes the quantitative visual-quality evidence unreliable; this must be corrected before the superiority claim can be assessed.","rationale":"Stress-testing the central claim requires checking whether the evidence for 'better-looking images' is solid. The reader correctly flags UIEBD's non-real reference images, but the paper acknowledges that limitation and partly mitigates it with no-reference metrics. The more direct problem is that the primary quantitative table contains an implausible value in the very row that is supposed to show superiority. A single outlier can be a typo, but without corrected numbers or variance estimates the paper's headline visual claim is not falsifiable from the published data. The downstream-task experiments are relevant, but they evaluate the coupled system (enhancement plus semantic collaborative learning plus task network), so they do not by themselves establish that the enhanced images are visually better. Thus the concrete check is to recompute the table from the released code. This is an internal-consistency concern, not a disagreement with community consensus, and it can be settled by computation. I therefore keep the reader's CONDITIONAL verdict unchanged: the issue is serious but correctable, and the central claim should not be accepted as stated until the table is verified.","tokens_in":19440,"tokens_out":11049,"duration_ms":108038,"concrete_test":"Download the released code and pretrained weights, rerun inference on the EUVP set, and compute UIQM with the authors' own metric implementation; specifically check whether 20.9151 reproduces. Then recompute the full Table 2 from the released model across three training seeds, reporting mean and standard deviation. If the EUVP UIQM falls to roughly 4.5 and the other HUPE margins shrink to within one standard deviation of SemiUIR/WaterFlow, the visual-superiority claim should be downgraded to 'competitive' rather than 'outstanding.'","verdict_should_be":"UNCHANGED","load_bearing_attack":"The visual-quality half of the central claim rests almost entirely on Table 2 (Section 4.3.2). In that table, the Ours row reports EUVP UIQM = 20.9151, while every other method on that dataset ranges from about 4.19 to 4.58, and the same model's UIQM on UIEBD, UCCS, and U45 is 4.20, 4.70, and 4.22, respectively. UIQM is a scalar image-quality index with no known scaling that would make 20.9 a plausible value for this model; the entry looks like a transcription error, possibly a PSNR-like value inserted into the UIQM column. Because the table also has no error bars or significance tests, and the HUPE row is mixed on other cells (e.g., it is below SemiUIR on UIEBD PSNR and below WaterFlow on EUVP UCIQE), a single corrupted cell makes it impossible to determine from the published data whether the claimed visual improvement is real or an artifact of reporting. If the entry is a typo, the corrected value may still leave some favorable cells, but the table as printed cannot support the abstract/conclusion claim that HUPE 'achieves good visual enhancement effects' and 'outperforms state-of-the-art methods.'","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes HUPE, an invertible network for underwater image enhancement that combines a reversible flow-based mapping (with Fourier-based affine coupling and a physical prior injector) with a semantic collaborative learning module that couples the enhancement network to downstream detection and segmentation networks. The method is evaluated on UIEBD, UCCS, U45, and EUVP for enhancement quality and on RUIE, Aquarium, and SUIM for downstream object detection and semantic segmentation, with ablations on loss components, heuristic prior inputs, the spatial-frequency block, the number of hybrid invertible blocks and flow steps, and the collaborative module. The central claim is that HUPE simultaneously improves visual quality and downstream perception performance.","tokens_in":1634,"tokens_out":1941,"duration_ms":58611,"significance":"If the empirical claims were fully supported, the paper would offer a useful architectural contribution: an invertible enhancement network with frequency-domain coupling, a physical-prior injector, and a mechanism for aligning enhancement features with task features. The open-source code, the breadth of datasets, and the inclusion of two downstream tasks are genuine strengths. However, the current quantitative evidence contains an apparent data anomaly and lacks error bars or significance tests, so the central superiority claim cannot yet be assessed reliably.","major_comments":[{"comment":"The EUVP UIQM entry for 'Ours' is 20.9151, while every other method on EUVP ranges from about 3.11 to 4.57 and the same model's UIQM on UIEBD, UCCS, and U45 is 4.20, 4.70, and 4.22, respectively. UIQM is a scalar image-quality measure with no known scaling that would make 20.9 a plausible value for this model, so this entry appears to be a transcription error. Because Table 2 is the main quantitative support for the abstract and conclusion claims of superior visual enhancement, and because neighboring cells already show HUPE below SemiUIR on UIEBD PSNR and below WaterFlow on EUVP UCIQE, the printed table cannot support the superiority claim until this cell is corrected and the full table is re-verified.","section":"Table 2, §4.3.2"},{"comment":"No error bars, confidence intervals, or significance tests are reported for any enhancement metric, and all comparisons appear to be single runs. Many observed differences are very small (for example, U45 UCIQE 0.6229 vs. 0.6241 and EUVP CEIQ 3.4679 vs. 3.4769), so without repeated trials or statistical tests they cannot be distinguished from run-to-run variation. This is load-bearing because the paper's central claim is a superiority claim over state-of-the-art methods.","section":"§4.3.2, Table 2"},{"comment":"The downstream comparison may not be apples-to-apples. HUPE's enhancement network and task perception module are jointly finetuned for an additional 20 epochs with the proposed semantic collaborative learning module, whereas the paper does not state whether the comparison enhancement methods receive an equivalent joint finetuning stage. If the task networks are trained on each enhanced image set with identical parameters but only HUPE gets the extra finetuning, the advantage could be due to that additional training stage rather than to the enhanced image quality. The protocol should be clarified, or an ablation should show that HUPE also wins without the joint finetuning advantage.","section":"§4.2, §4.4, §4.5.6"},{"comment":"The paper correctly acknowledges that UIEBD reference images are outputs of other underwater image enhancement methods rather than real captures of clear underwater scenes. This means the supervised PSNR/SSIM numbers and the supervised training objective measure imitation of those particular enhancers, not physical fidelity to true in-air scenes. The visual-quality claim therefore depends heavily on the non-reference metrics and qualitative comparisons; with the EUVP UIQM anomaly and the absence of statistical tests, that support is currently insufficient. A human perceptual study or a validation on data with real ground-truth clear images would substantially strengthen the claim.","section":"§4.3.2"}],"minor_comments":[{"comment":"The sentence 'RUIE and Aquarium ... were used to evaluate the effectiveness of Tatarflow' contains a typo; it should refer to HUPE (or WaterFlow). The paper should also disambiguate 'WaterFlow' as both the preliminary version name and a comparison baseline.","section":"§4.2"},{"comment":"The text for the Hybrid Invertible Block ablation refers to panel (c) of Fig. 14, and the flow-step ablation refers to panel (d), but the caption assigns (d) to the number of HIBs and (e) to the number of flow steps. Please renumber the cross-references.","section":"§4.5.4, §4.5.5, Fig. 14"},{"comment":"In Eq. (3), the notation phi_i(u_i^1,F) is used for both the multiplicative and additive terms of the affine coupling; two different functions (scale and shift) should be defined. In Eq. (4), the symbol k is used as the imaginary unit after i was already used as an index; please use a consistent and clearly defined imaginary-unit symbol.","section":"Eqs. (3), (4)"},{"comment":"The transmission model is written as J^c = I^c/t + B^c(t-1)/t, which differs from the standard form J = (I-B)/t + B. Please check the algebra, since the heuristic prior injector is derived from this equation.","section":"Eq. (1)"},{"comment":"The ablation figure caption refers to 'L1, Ls and Lc' as 'L1 loss, style loss, and contrasive loss,' while the loss equations define L_f as a frequency loss; the terminology should be unified so that the style/frequency loss has one name.","section":"Fig. 13 and §3.4"}],"recommendation":"major_revision","confidential_remarks":"The paper is an extended version of the authors' prior WaterFlow work, and the incremental value of the frequency-aware coupling and semantic collaborative learning should be made explicit through direct comparisons and a clearer novelty discussion. Given the apparent outlier in Table 2 and the missing statistical support, I would request an updated table, repeated runs with variance reporting, and a clarification of the downstream training protocol before considering the paper for publication."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: this is a decent incremental method paper for underwater image enhancement, but the main quantitative claim is currently unsupported because Table 2 contains an obvious error. The HUPE row lists EUVP UIQM as 20.9151 while every other UIQM on that dataset is between about 3.1 and 4.6, and the same model's UIQM elsewhere is 4.2–4.7. That is a typo or a wrong-value paste, not a real measurement. As printed, the table cannot be used to conclude HUPE is superior on visual quality.\n\nWhat's genuinely new: the paper extends the authors' WaterFlow with a frequency-aware affine coupling layer and, more importantly, a semantic collaborative learning module that feeds task-aware features into the enhancement network. The downstream experiments on detection and segmentation, plus the ablation of the SCL module, are the most interesting part. The authors also provide code, which is good practice.\n\nThe soft spots beyond the outlier: no error bars or significance tests anywhere; all quantitative results appear to come from single runs. With the outlier removed, the method is still competitive but not clearly state-of-the-art on several cells. The paper is honest about the UIEBD references being outputs of other enhancement methods rather than real clean underwater captures, but it then still uses PSNR/SSIM against those references as evidence. That is a genuine caveat for interpreting the supervised numbers. Finally, the architecture of the heuristic prior encoder is deferred to the previous WaterFlow paper, so the method is not fully self-contained.\n\nWho this is for: people working on underwater enhancement and on joint low-level/high-level optimization. The SCL idea is worth knowing about, and the code makes it easy to try.\n\nMy recommendation: yes, send it to peer review. It deserves a serious referee. But the authors need to fix the Table 2 error, report variance or multiple runs, and ideally discuss the reference-imitation problem more carefully. I'd expect a conditional accept after major revision.","headline":"Solid incremental method paper whose central quantitative claim is undercut by a glaring Table 2 error (EUVP UIQM = 20.9151); the semantic collaborative learning idea is the most interesting part.","tokens_in":20253,"tokens_out":2724,"would_cite":false,"duration_ms":25382,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An invertible network with Fourier and semantic guidance improves underwater image quality and downstream perception tasks in one pass.","keywords":["underwater image enhancement","invertible network","normalizing flow","Fourier transform","semantic segmentation","object detection","heuristic prior","collaborative learning"],"falsifier":"Take a set of underwater scenes for which true clear-water reference photographs exist (for example, the same objects photographed in a laboratory tank with clear water), train and test HUPE on those pairs, and compare PSNR/SSIM and downstream detection/segmentation accuracy against the baselines; if HUPE does not outperform them on these physically grounded references, its claim of perceptual enhancement is not supported.","tokens_in":19216,"feed_emoji":"🌊","tokens_out":5160,"duration_ms":38799,"temperature":0.7,"pith_summary":"This paper claims that underwater image enhancement can serve both visual quality and downstream perception tasks at once by building the enhancement as an invertible mapping with physically motivated priors and a semantic collaboration mechanism. The authors introduce HUPE, a network that learns a bidirectional translation between degraded underwater images and clear counterparts, while a feature-level collaboration module feeds semantic cues from detection or segmentation networks back into the enhancement process. They report that HUPE produces higher-quality enhanced images than state-of-the-art methods on standard benchmarks and, more importantly, yields better object detection and semantic segmentation results when the enhanced images are used as inputs. If correct, this would make enhancement a more practical preprocessing tool for underwater robotics and monitoring.","feed_headline":"Invertible network lifts underwater images and downstream perception","feed_subtitle":"Reversible mapping with Fourier and semantic guidance beats prior enhancers on quality, detection, and segmentation.","key_machinery":"The central object is the Hybrid Invertible Block (HIB), a stack of reversible transformations (Actnorm, invertible 1x1 convolution, a Heuristic Prior Injector, and a Frequency-Aware Affine Coupling) that maps an underwater image to a clear image in the forward pass and back again in the reverse pass. The heuristic prior injector estimates ambient light and transmission using depth and gradient maps derived from a dark-channel prior; the frequency-aware coupling uses the Fourier transform so the network sees both phase (semantic) and amplitude (style) information. On top of this, the Semantic Collaborative Learning module, with a Meta-Feature Generator and Feature Transformation Block, inserts task-aware features into the enhancement network's intermediate layers during training, so the enhancer learns to keep semantic structure. The whole system is trained with a guide loss, an enhancement loss combining contrastive, frequency, and bilateral terms, and a task loss.","core_discovery":"On the paper's own terms, the discovery is that a heuristic invertible network with an embedded Fourier transform and semantic collaborative learning can simultaneously improve the visual appearance of underwater images and make them more useful for object detection and semantic segmentation. The authors argue that the reversibility of the mapping preserves information that ordinary one-way enhancement networks lose, the Fourier-based affine coupling captures relationships in both spatial and frequency domains, and the collaborative learning module narrows the feature gap between low-level enhancement and high-level perception. They support this with quantitative and qualitative experiments on UIEBD, UCCS, U45, and EUVP for enhancement, and on RUIE, Aquarium, and SUIM for detection and segmentation, concluding that HUPE is both a strong enhancer and a strong preprocessor for perception tasks.","pith_inferences":["The method's reliance on synthetic or algorithm-generated references means the reported PSNR/SSIM gains may overstate true physical restoration; a fairer test would use images with known clear-water counterparts or human preference studies.","The semantic collaboration idea could be transferred to other degradation domains, such as low-light or haze, where the same tension between visual quality and perception-task performance exists.","Because the invertible network is a normalizing-flow-style architecture, the learned mapping could potentially be used for uncertainty estimation or generation of diverse plausible clear images, though the paper does not explore this."],"forward_implications":["Enhanced images from HUPE can be fed directly into standard detectors and segmenters, improving their accuracy without retraining the perception networks' architecture.","Because the mapping is invertible, the same network can be used to reconstruct the original underwater image from an enhanced one, which may help downstream tasks that need consistency or reversibility.","The joint optimization scheme suggests a general recipe: any low-level restoration network can be coupled with a downstream perception network through meta-features to produce task-oriented restorations.","The heuristic priors (depth and gradient) reduce the network's reliance on large paired datasets, making the approach more adaptable to new underwater conditions."],"supporting_citations":[{"why":"Supplies the prior heuristic-encoder architecture and the WaterFlow baseline this work extends with Fourier coupling and semantic collaboration.","marker":"Zhang et al., 2023"},{"why":"Provides the invertible building blocks (Actnorm and 1x1 convolution) that the Hybrid Invertible Block is composed of.","marker":"Kingma and Dhariwal, 2018"},{"why":"Provides the UIEBD benchmark whose reference images are used for training, supervised metrics, and the paired comparison baseline.","marker":"Li et al., 2019"},{"why":"Formulates the underwater image formation model that the heuristic prior injector uses to estimate ambient light and transmission.","marker":"Chiang and Chen, 2011"},{"why":"Used for transmission estimation in underwater images, informing the depth and transmission priors.","marker":"Drews et al., 2013"},{"why":"Generalizes the dark channel prior to estimate depth and gradient maps for the heuristic prior generator.","marker":"Peng et al., 2018"},{"why":"Defines focal loss, used as the classification loss in the detection task loss.","marker":"Lin et al., 2017"},{"why":"Provides the VGG19 network whose features are used in the contrastive enhancement loss.","marker":"Simonyan and Zisserman, 2014"}],"fun_headline_variants":["HUPE: invertible net enhances underwater images and perception","Invertible Fourier net improves underwater visuals and AI tasks","Underwater enhancement meets detection: HUPE's dual boost","HUPE: reversible mapping sharpens underwater scenes for AI","Semantic-collaborative invertible network lifts underwater perception"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The training and evaluation rely on paired 'clear' reference images that are themselves outputs of other underwater enhancement methods rather than true photographs of the same scenes in clear water, so the supervised quality scores may reward imitating other enhancers instead of true physical restoration.","fun_headline_variants_meta":{"raw":{"variants":["HUPE: invertible net enhances underwater images and perception","Invertible Fourier net improves underwater visuals and AI tasks","Underwater enhancement meets detection: HUPE's dual boost","HUPE: reversible mapping sharpens underwater scenes for AI","Semantic-collaborative invertible network lifts underwater perception"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000582,"raw_usage":{"total_tokens":2711,"prompt_tokens":885,"completion_tokens":1826,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":501,"completion_tokens_details":{"reasoning_tokens":1755}},"tokens_in":501,"tokens_out":1826,"duration_ms":17105,"temperature":1.0,"reasoning_tokens":1755,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-12T11:19:19.553948+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Take a set of underwater scenes for which true clear-water reference photographs exist (for example, the same objects photographed in a laboratory tank with clear water), train and test HUPE on those pairs, and compare PSNR/SSIM and downstream detection/segmentation accuracy against the baselines; if HUPE does not outperform them on these physically grounded references, its claim of perceptual enhancement is not supported.","supporting_citations":[{"cited_title":", Goyal , P","cited_arxiv_id":null,"evidence_quote":"Defines focal loss, used as the classification loss in the detection task loss."},{"cited_title":", Cao , K","cited_arxiv_id":null,"evidence_quote":"Generalizes the dark channel prior to estimate depth and gradient maps for the heuristic prior generator."},{"cited_title":", Jiang , Z","cited_arxiv_id":null,"evidence_quote":"Supplies the prior heuristic-encoder architecture and the WaterFlow baseline this work extends with Fourier coupling and semantic collaboration."},{"cited_title":", Dhariwal , P","cited_arxiv_id":null,"evidence_quote":"Provides the invertible building blocks (Actnorm and 1x1 convolution) that the Hybrid Invertible Block is composed of."},{"cited_title":", Chen , Y.-C","cited_arxiv_id":null,"evidence_quote":"Formulates the underwater image formation model that the heuristic prior injector uses to estimate ambient light and transmission."},{"cited_title":", Nascimento , E","cited_arxiv_id":null,"evidence_quote":"Used for transmission estimation in underwater images, informing the depth and transmission priors."}],"review_version":1}