{"id":"73a0d01c-b95e-4f83-8dcc-75a2a322e7f4","arxiv_id":"2608.08661","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":5,"one_line_summary":"PROTEUS couples spatially varying degradation cues with a task-oriented latent code to restore underwater images, reaching top PSNR and LPIPS on U90 and LSUI-400 with 2.61M parameters.","lead":"PROTEUS is a new deep-learning network for underwater image restoration that uses degradation cues to guide feature processing and a learned control code to regulate skip connections. It reports state-of-the-art or competitive results on five paired and four non-reference underwater benchmarks, using only 2.61M parameters.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The reported benchmark margins on U90 and LSUI are small enough that evaluation-protocol differences across methods could change the claimed ranking.","rationale":"The stress-test pass agrees with the reader's weakest-assumption analysis: the benchmark tables are not produced under a unified evaluation protocol, and the performance gaps are small enough that protocol differences could plausibly change the ranking. I examined the architectural claims, the ablation tables, and the supplementary diagnostics for internal inconsistencies or other load-bearing weaknesses. The ablation results are qualitatively consistent with the stated design choices, and the gate-intervention and representation diagnostics are thoughtful and appropriately cautious. The one additional observation is the SSIM discrepancy between Table 1 (0.933) and Table S3 (0.935) for the same U90 evaluation; this does not by itself overturn the paper, but it underscores that even PROTEUS's own reported numbers are not perfectly reproducible from the text. Since the central claim depends on comparing PROTEUS's numbers with numbers taken from other papers, the correct verdict remains CONDITIONAL, which is the reader's verdict; no change is needed.","tokens_in":18467,"tokens_out":4158,"duration_ms":47002,"concrete_test":"Release the PROTEUS evaluation script and checkpoint, obtain or re-run the official code and checkpoints for the closest competitors (at least UniUIR, MoCE-IR, AdaIR, and WWE-UIE), and evaluate all methods under the exact supplementary protocol: resize inputs and references to 256x256, use the same U90 and LSUI-400 held-out splits, and compute PSNR, SSIM, and AlexNet LPIPS from the same codebase. If the resulting PSNR/SSIM/LPIPS re-ranking no longer shows PROTEUS best on U90 or LSUI-400, or if the U90 SSIM value still disagrees with Table S3's 0.935, the headline performance claim should be revised.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim is that PROTEUS achieves the best or joint-best restoration quality on U90 and LSUI-400 with a favorable cost balance. For that claim to hold, the comparison must be fair across methods. The supplement states that the main tables combine previously published baseline numbers with PROTEUS results evaluated from the authors' own checkpoints, and explicitly calls these 'benchmark-level comparisons rather than jointly retrained results under a unified codebase.' This is the least secure link: the margins are narrow, 0.39 dB over UniUIR and 0.57 dB over MoCE-IR on U90, and 0.57 dB over UniUIR on LSUI-400. Differences in input resizing, test-crop strategy, LPIPS backbone, evaluation resolution, or the exact test split version can shift PSNR by more than these gaps. Because PROTEUS's own numbers are computed at 256x256 resolution with AlexNet LPIPS, any baseline published under a different protocol is not strictly comparable. An internal discrepancy reinforces the concern: main Table 1 reports U90 SSIM 0.933 for PROTEUS, while Table S3 reports 0.935 for the 'Full spatial guide' row, which is the same checkpoint evaluated on the same 90 U90 pairs. Either one table is mis-rounded or the evaluation conditions differ between them; in either case, a unified re-evaluation is needed before the ranking claim can be taken as established. The architectural components are plausible and the ablation trends are internally consistent, but the central quantitative claim is conditional on the comparison protocol being fair.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"PROTEUS is a U-Net-based underwater image restoration method that couples two mechanisms: a Guided Dynamic Feature Modulation Block (GDFMB), which uses a Gray-Edge-derived guide image and deformable cross-attention with Gaussian basis expansion to adapt feature processing across stages, and a task-oriented latent controller that maps the bottleneck feature into a control code F_z, regularizes it with alignment, orthogonality, and contrastive losses, and uses it to modulate skip-feature channels through Attention Gates. The paper claims highly competitive restoration performance on five paired benchmarks (U90, LSUI, UFO-120, EUVP-Scene, EUVP-Dark) and four non-reference benchmarks, with a favorable cost profile of 2.61M parameters and 18.52G FLOPs. The evidence includes ablations, same-checkpoint test-time interventions, bootstrap confidence intervals, cross-dataset transfer tests, and explicit representation diagnostics that avoid overclaiming a clean-latent interpretation of F_z.","tokens_in":18765,"tokens_out":5957,"duration_ms":67942,"significance":"If the headline results hold under a unified evaluation protocol, PROTEUS would offer a strong quality/cost trade-off and a reusable idea: using a structured bottleneck code to regulate skip-feature reuse rather than merely to condition the decoder. The manuscript is more careful than many papers in this area: the same-checkpoint interventions with paired bootstrap CIs, per-image gain distributions, real-degradation stratification, and the explicit statement that F_z is not a metrically cleaner embedding all provide genuinely useful evidence for the architectural claims. However, the central quantitative claim of state-of-the-art ranking currently depends on a benchmark-level comparison that mixes previously published numbers with the authors' own checkpoints, and there is an internal inconsistency between the main tables and the supplementary tables. These issues are fixable but must be addressed before the ranking claims can be considered established.","major_comments":[{"comment":"The headline claim that PROTEUS achieves the best PSNR and LPIPS on U90 and LSUI-400 is not yet supported by a controlled comparison. The supplement states that the benchmark tables 'combine results reported by the cited methods with PROTEUS results evaluated using our checkpoints' and should be viewed as 'benchmark-level comparisons rather than jointly retrained results under a unified codebase.' Given the small margins (0.39 dB over UniUIR on U90 and 0.57 dB over UniUIR on LSUI-400), differences in input resizing, test-crop strategy, LPIPS backbone version, or test split can plausibly change the ranking. The authors should either re-evaluate the closest baselines under the same protocol used for PROTEUS (resize to 256×256, AlexNet LPIPS, identical split) or explicitly reframe the claims as benchmark-level, with all ranking statements made conditional on that caveat.","section":"Supplementary, 'Evaluation protocol and result provenance'; Tables 1–2"},{"comment":"There is an internal inconsistency in the reported full-model metrics. Main Table 1 lists PROTEUS on U90 as PSNR 25.50, SSIM 0.933, LPIPS 0.081 and on LSUI-400 as PSNR 28.99, SSIM 0.913, LPIPS 0.083. Table S3, described as the same checkpoint evaluated on all 90 U90 pairs and 400 LSUI pairs, gives U90 SSIM 0.935 and LSUI SSIM 0.921 for the 'Full spatial guide' row, with identical PSNR and LPIPS values. Since the PSNR and LPIPS match exactly, simple rounding cannot explain both SSIM discrepancies (0.933 vs. 0.935 and 0.913 vs. 0.921). The authors must reconcile these numbers, state explicitly how unrounded values are handled, and ensure that every table reporting the same model reports the same metric values, otherwise the reliability of the quantitative results cannot be audited.","section":"Table 1 vs. Table S3"},{"comment":"The paper should more prominently flag that the latent-loss weight sensitivity study is a fixed-coordinate local adaptation experiment, not an ablation of the trained model. The supplement correctly says 'adapted scores are not comparable to the unadapted 25.50 dB main-paper result,' but the main text's representation-and-gate diagnostics paragraph cites the three-seed analysis without this caveat. This matters because the same section reports that increasing the alignment weight reduces but never eliminates the F_z-vs-raw-bottleneck deficit, which is an important negative result for the clean-latent interpretation. The main text should carry the same qualification so readers do not mistake the local-sensitivity numbers for main-model properties.","section":"Supplementary, Table S5 and 'Latent-Loss Weight Sensitivity'"}],"minor_comments":[{"comment":"The loss-removal rows in Table S2 are independently trained runs, not seed-paired interventions, and the supplement already describes them as 'descriptive ablation evidence rather than paired statistical effects.' The main text should repeat this caveat in the paragraph that cites the loss ablations, rather than presenting the 25.50-to-25.24 drop as a precisely measured effect.","section":"Supplementary, Table S2 and 'Additional Ablation Results'"},{"comment":"The reference list contains duplicated entries: Islam et al. 2020a and 2020b refer to the same paper, and Zhang et al. 2022a and 2022b also appear to be the same publication. These duplicates should be merged and cited with distinct keys only where different works are intended.","section":"References"},{"comment":"Figure 2 is dense and has a very long caption; in the final version the module names, data-flow arrows, and augmentation branches should be redrawn at larger size and with legible labels, because several components (GDAB, GBE, Attention Gate, photometric branch) are difficult to resolve from the current rendering.","section":"Figure 2"},{"comment":"The FLOPs and parameter counts for competing methods are taken from their respective papers, and the supplement does not state whether the same FLOPs-counting conventions (e.g., whether batch dimensions or the final convolution are included) apply to all entries. A short footnote stating the counting tool and input size for all rows would prevent apples-to-oranges efficiency comparisons.","section":"Tables 1–2, FLOPs reporting"}],"recommendation":"major_revision","confidential_remarks":"The paper is within scope and the architectural idea is worth considering, but the central comparative claims cannot be accepted as they stand because the benchmark tables mix protocols and because of the Table 1 versus Table S3 inconsistency. I would require a re-evaluation of at least the closest baselines (UniUIR, MoCE-IR, WWE-UIE) under a common protocol, or a clear reframing of all ranking claims as benchmark-level. The same-checkpoint intervention analyses and the honest treatment of the latent-representation diagnostics are strong assets and should be emphasized."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"You should know this is a genuinely solid piece of engineering. PROTEUS couples degradation-guided feature modulation with a task-oriented latent control code that gates skip connections. The individual parts are familiar, but the specific combination, plus the systematic evaluation, is new enough to matter in the underwater imaging subfield.\n\nWhat impressed me: the paper does not oversell its latent code. It explicitly reports that the control code is farther from the ground-truth latent than the raw bottleneck, and it runs same-checkpoint interventions to isolate the contribution of spatial guidance and gate conditioning. The bootstrap confidence intervals and per-image gain analyses are a cut above what you usually see in this literature. The ablations are internally consistent, and the supplementary is candid about the limits of the benchmark-level comparison.\n\nNow the soft spots. The headline ranking on U90 and LSUI is not fully verifiable because the main tables mix previously published baseline numbers with PROTEUS's own checkpoints. The margins over the next best methods are 0.4–0.6 dB, and protocol differences in resizing, crop strategy, LPIPS backbone, or test split can easily shift PSNR by that amount. The supplementary acknowledges this, but the paper still states \"best\" results as if the comparison were unified. There is also a concrete internal discrepancy: main Table 1 reports U90 SSIM 0.933 for PROTEUS, while Table S3 reports 0.935 for the same checkpoint and same 90 images. That needs to be explained or corrected, because it undermines confidence in the exact numbers. No code or checkpoints are released, so I cannot at present reproduce any of it.\n\nThe central architecture is plausible and the ablation logic is sound. The \"task-oriented\" framing is operational rather than theoretical, but the paper consistently stays within that framing, which I respect. The GEPM preprocessing is shown to be nearly interchangeable with feeding the raw input as the guide, so the paper wisely de-emphasizes that particular claim.\n\nFor peer review: yes, send it out. The experimental effort and the honesty of the diagnostics deserve serious referee time. The authors should be asked to release code, fix the SSIM discrepancy, and either provide a unified re-evaluation of baselines or tone down the ranking claims. It is a useful contribution to the underwater restoration niche, not a field-reshaping one.\n\nI would not cite it yet, because the central quantitative claim is unverifiable without code and the benchmark numbers are not unified. If the code appears and the discrepancy is resolved, it becomes a standard citation for the subfield.","headline":"A careful, transparently-reported underwater restoration paper whose main SOTA claim rests on mixed-protocol benchmarks and a small internal inconsistency; worth refereeing, but the authors should release code and confirm the ranking under a unified evaluation.","tokens_in":19301,"tokens_out":1717,"would_cite":false,"duration_ms":21448,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper claims that underwater restoration improves when degradation cues guide feature adaptation while a learned latent code regulates skip connections, achieving the best PSNR and LPIPS on U90 and LSUI-400 with a 2.61M-parameter…","keywords":["underwater image restoration","degradation-guided feature modulation","task-oriented latent control","skip connection regulation","U-Net","image enhancement","deep learning","underwater benchmarks"],"falsifier":"Retrain every compared method and PROTEUS from scratch in one codebase with identical UIEB splits, patch size, augmentation, and training epochs, then recompute U90 and LSUI-400 PSNR, SSIM, and LPIPS: if PROTEUS no longer holds the best PSNR and LPIPS, the claimed advantage is a setup artefact.","tokens_in":18268,"feed_emoji":"🌊","tokens_out":10928,"duration_ms":107177,"temperature":0.7,"pith_summary":"Underwater images carry degradation information that behaves two ways: it signals what needs correcting, and it contaminates the features a decoder reuses. PROTEUS is built on the claim that a restoration network should exploit both. Its degradation-guided feature modulation block steers feature processing with spatially varying cues, while a task-oriented latent controller learns a bottleneck code that decides which skip-connection channels to attenuate during decoding. On paired benchmarks the paper reports the best PSNR and LPIPS on U90 (25.50 dB and 0.081) and LSUI-400 (28.99 dB and 0.083), with a small footprint of 2.61M parameters and 18.52G FLOPs.","feed_headline":"A 2.6M-parameter net leads underwater restoration benchmarks","feed_subtitle":"Degradation cues guide feature adaptation; a learned latent code regulates skip channels for better color at low compute.","key_machinery":"The central machinery is the pair formed by the Guided Dynamic Feature Modulation Block (GDFMB) and the task-oriented latent controller. GDFMB uses a guide image derived from a Gray-Edge colour-bias prior, samples spatially varying degradation cues through deformable cross-attention, expands them through Gaussian basis functions with a selective fusion gate, and adds a Gradient Fusion Block to preserve edges. The latent controller maps the bottleneck feature into a control code $F_z$ and a complementary residual $F_{res}$, regularises $F_z$ with reference alignment, orthogonality, and contrastive loss, and uses it to generate channel-wise Attention Gate weights that rescale skip features before decoder fusion. Together they turn degradation information into two separated functions: guiding where features are transformed and deciding which skip channels to pass through, without assuming either representation is a physically clean image.","core_discovery":"PROTEUS claims that underwater degradation information should be treated as a dual-use signal rather than something to simply estimate and remove. Spatially varying degradation cues from the input are used to guide adaptive feature transformation at every encoder–decoder stage, while degradation-entangled encoder features are regulated before reuse by a bottleneck control code applied channel-wise to skip connections. The control code is explicitly task-oriented, not a clean-image embedding: it is regularised by alignment, orthogonality, and contrastive separation, and diagnostics reported by the paper show it is actually farther from the ground-truth latent than the unfiltered bottleneck feature. By coupling these two controls, the paper obtains best PSNR and LPIPS on U90 and LSUI-400, joint-best SSIM on U90, and competitive results across four non-reference benchmarks, at 2.61M parameters and 18.52G FLOPs.","pith_inferences":["Because the paper finds the raw underwater input nearly matches the preprocessed Gray-Edge guide in guiding the network, a testable extension is to replace the hand-designed guide with a learned adaptable preprocessing module and measure whether spatial guidance gains grow.","The task-oriented control code, which is explicitly not a clean-latent projection, may transfer to other ill-posed restoration settings such as low-light or scattering-medium imaging where clean/degraded separation is not well defined.","The gate diagnostics showing stable channel modulation with weak per-channel degradation correlation suggest the controller could be exposed as a user-adjustable colour and contrast control in practical underwater systems, a role the paper does not claim.","Per-image and cross-dataset interventions show spatial guidance helps most but not all images, so a confidence or severity estimator that decides when to disable spatial guidance is a natural follow-up."],"forward_implications":["If PROTEUS is right, skip connections in underwater restoration networks should be seen as a regulated information pathway rather than a pass-through of encoder features.","The two controls are complementary: ablations show that removing either the latent-conditioned Attention Gate or the deformable spatial guidance lowers PSNR, and visual analyses show combining them gives the best colour and detail.","The reported representation diagnostics indicate that a useful control code need not be closer to a clean-image latent, which supports task-oriented regularisation for restoration pipelines that cannot assume a clean/degraded decomposition.","The claimed gains come with a small footprint of 2.61M parameters and 18.52G FLOPs, so the dual-control strategy is a plausible candidate for lightweight underwater imaging systems."],"supporting_citations":[{"why":"Supplies the barycentric latent-geometry idea that PROTEUS adapts into its task-oriented latent regularisation, and provides precomputed negative latents used in the contrastive term.","marker":"(Tang et al. 2025)"},{"why":"Provides the five standard restoration losses (Charbonnier, HVI colour, SSIM, VGG, edge) that make up the PROTEUS objective, and is the closest baseline on U90 and LSUI.","marker":"(Cheng et al. 2026)"},{"why":"Supplies the UIEB training set and U90 benchmark test set used for the main quantitative comparison.","marker":"(Li et al. 2020)"},{"why":"Provides the LSUI dataset and the Ushape transformer baseline; LSUI-400 is one of the two paired benchmarks where PROTEUS reports the best PSNR and LPIPS.","marker":"(Peng, Zhu, and Bian 2023)"},{"why":"The Gray-Edge prior is the basis of the guide image generated by the Gray-Edge Prior Module that conditions GDFMB.","marker":"(Van De Weijer, Gevers, and Gijsenij 2007)"},{"why":"FUnIE is a source of negative latents for the contrastive regularisation of the control code.","marker":"(Islam, Xia, and Sattar 2020a)"}],"fun_headline_variants":["Degradation guides features and a task code regulates skip connections underwater","Underwater restoration with degradation-guided features and task-oriented latent control","2.6M-param PROTEUS uses degradation for features and task code for skip control","PROTEUS: degradation guides feature adaptation and a task-tuned code controls skips"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The headline comparison assumes that published benchmark numbers for competing methods can be fairly compared with PROTEUS numbers produced from the authors' own checkpoints; the paper explicitly calls these benchmark-level comparisons rather than jointly retrained results under a unified codebase, so if training protocols, data splits, or evaluation settings differ across methods, the reported gaps may be setup differences rather than architectural superiority.","fun_headline_variants_meta":{"raw":{"variants":["Degradation guides features and a task code regulates skip connections underwater","Underwater restoration with degradation-guided features and task-oriented latent control","2.6M-param PROTEUS uses degradation for features and task code for skip control","PROTEUS: degradation guides feature adaptation and a task-tuned code controls skips"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000709,"raw_usage":{"total_tokens":3169,"prompt_tokens":896,"completion_tokens":2273,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":512,"completion_tokens_details":{"reasoning_tokens":2190}},"tokens_in":512,"tokens_out":2273,"duration_ms":19032,"temperature":1.0,"reasoning_tokens":2190,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:27:29.246838+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Retrain every compared method and PROTEUS from scratch in one codebase with identical UIEB splits, patch size, augmentation, and training epochs, then recompute U90 and LSUI-400 PSNR, SSIM, and LPIPS: if PROTEUS no longer holds the best PSNR and LPIPS, the claimed advantage is a setup artefact.","supporting_citations":[],"review_version":1}