{"id":"95ffcf5c-5da1-45d9-935b-122d606d600d","arxiv_id":"2607.24002","paper_version":1,"verdict":"REJECT","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"high","formal_verification":"none","parameter_count":6,"one_line_summary":"A Fourier-amplitude-guided multi-scale attention network is claimed to outperform prior low-light enhancement methods by wide margins on LOL, SID, SMID, and SDSD benchmarks.","lead":"This paper presents MSFT, a supervised neural network for low-light image enhancement that combines Fourier-transform amplitude guidance with multi-scale attention modules. The authors claim large benchmark gains over prior methods, but the submitted manuscript contains internally inconsistent comparison tables and no code, so the headline results are not independently verifiable.","discovery_kind":"new_method","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table-level inconsistencies (Retinexformer 25.16 vs 22.43; duplicate W/O MSA/W/O DRA rows; final scores equal 12-run means) leave the claimed SOTA margins unverified until baselines are rerun under a common protocol.","rationale":"The reader's weakest assumption—that Table 4 baseline numbers are accurate and protocol-comparable—is exactly where the central claim fails. I agree, but I find additional internal indicators that the tables are unreliable: the Retinexformer LOL-v1 value changes between Table 4 and Table 5; the W/O MSA and W/O DRA ablation rows are nearly identical despite one being a submodule of the other; and the reported final scores coincide exactly with 12-run means despite a stated best-of-12 selection. These are concrete, locatable inconsistencies that undermine the empirical comparison before any external reproduction is attempted. The architecture itself is a plausible incremental combination of Fourier-domain guidance and multi-scale attention, but the manuscript's burden is to demonstrate SOTA performance with trustworthy numbers. That burden is not met. The correct disposition remains REJECT, pending release of code and corrected, reproducible tables.","tokens_in":22082,"tokens_out":7679,"duration_ms":70286,"concrete_test":"Run the official Retinexformer implementation on LOL-v1 and SDSD-outdoor using the exact train/test split and preprocessing described in §4.1, and compare the resulting PSNR to 25.16 (Table 4), 22.43 (Table 5), and 29.84 (Table 4). If the rerun differs from the relevant table entry by more than ~0.5 dB, the claimed margins are not established. Ideally also run MSFT from released code under the same protocol and check whether 41.76 is a selected-max or a mean.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central claim is its benchmark superiority: 41.76 dB PSNR on SDSD-outdoor, 11.92 dB over Retinexformer. This is an empirical claim, and its only evidence is Table 4. That table is not internally reliable. (i) Table 4 gives Retinexformer LOL-v1 PSNR 25.16 dB; Table 5 gives 22.43 dB for the same dataset, and 22.43 is exactly Restormer's Table 4 value—a transcription collision. (ii) Table 7's W/O MSA row and Table 8's W/O DRA row are numerically identical to 2 decimals on PSNR and near-identical on SSIM (26.94/26.94, SSIM 0.837/0.837 on LOL-v1), yet DRA is a component of MSA; removing MSA should not equal removing only DRA. (iii) The final MSFT values in Table 4 equal the 12-run means in Table 3 on all seven datasets, although §4.1 states the run with highest validation PSNR was chosen; with nonzero standard deviations, best-of-12 should be at or above the mean, not exactly equal. These are locatable, mechanical inconsistencies. No code, weights, or baseline reproduction protocol is provided. If Table 4 baselines cannot be trusted, the 11.92 dB advantage and 'outperforms all' conclusion have no verified foundation.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes MSFT, a supervised low-light image enhancement network with a U-shaped architecture combining Fourier-transform-guided multi-scale attention (FTG-MSA), a Multi-Shape Synergistic Attention (MSSA) module built from SCSA, MSA (DSA/DRA), and CMUNeXt, plus multi-scale feature alignment. The central claim is empirical: MSFT outperforms state-of-the-art methods on seven benchmarks (LOL-v1, LOL-v2-real/syn, SID, SMID, SDSD-indoor/outdoor), with highlighted gains of 11.92 dB PSNR over Retinexformer on SDSD-outdoor (41.76 dB, SSIM 0.988). The paper also reports ablations, complexity statistics, and repeated-run statistics.","tokens_in":22414,"tokens_out":5620,"duration_ms":52247,"significance":"If the reported results are accurate, the claimed margins—particularly the 9–16 dB gains on SDSD datasets—would be a substantial advance in low-light image enhancement. The architecture description is fairly detailed, and the attempt to report run-to-run variability is a good practice. However, the paper provides no code, trained models, or baseline reproduction protocol, and several core tables contain mechanical inconsistencies. Because the headline contribution is an empirical benchmark comparison, these issues directly affect the validity of the stated results. The strengths are the clear modular design and the inclusion of repeated-run statistics; the weakness is that the numerical evidence is not internally reliable.","major_comments":[{"comment":"In Table 4, Retinexformer's LOL-v1 PSNR is 25.16 dB; in Table 5, the same model is reported as 22.43 dB on the same dataset. 22.43 dB is exactly Restormer's LOL-v1 value in Table 4. This collision indicates that at least one baseline number is wrong. Since the abstract's 11.92 dB margin on SDSD-outdoor and the 'outperforms Retinexformer' claim rest on the Table 4 baseline values, this is a load-bearing error.","section":"Table 4 vs Table 5, §4.2"},{"comment":"The 'W/O MSA' row in Table 7 is numerically identical in PSNR on all seven datasets (26.94, 25.95, 25.55, 26.54, 29.80, 39.04, 41.75) to the 'W/O DRA' row in Table 8, with SSIM equal except for one digit on SID. Since MSA is described as a parallel combination of DSA and DRA, removing the whole MSA should differ from removing only the DRA branch. The duplication suggests one row was copied from the other, so the ablation study does not support the claimed attribution of gains to DRA or MSA.","section":"Tables 7 and 8, §4.3"},{"comment":"The MSFT values in Table 4 are exactly the 12-run means reported in Table 3 on all seven datasets (e.g., 27.48, 25.78, 41.76). Yet §4.1 states that the model with the highest validation PSNR was chosen from the 12 runs. With reported standard deviations of 0.44–1.41 dB, the best-of-12 value cannot equal the mean to two decimals on every dataset. Table 4 therefore appears to report mean performance, which contradicts the stated selection protocol and makes the comparison against baselines potentially unfair.","section":"§4.1, Tables 3 and 4"},{"comment":"The categories in Table 4 are mislabeled: SNR-Net and Retinexformer are listed in the 'UL' (unsupervised) group, and the text explicitly calls them 'unsupervised methods' in §4.2. Both are supervised methods trained on paired data. This misclassification obscures the comparison structure and raises doubt about the care with which the baseline numbers were assembled.","section":"Table 4, §4.2"},{"comment":"No code, trained weights, or baseline reproduction details are provided, and the baseline numbers are not sourced beyond the table. Given the internal inconsistencies listed above, the central claim that MSFT outperforms all SOTA methods is not supported by the submitted evidence. The authors should make the evaluation protocol public and rerun or verify all baselines under a common training setting.","section":"Overall empirical evidence"}],"minor_comments":[{"comment":"There are several typos: 'priori channels' should be 'prior channels'; the Figure 1 caption contains 'of of'; and 'superivised' should be 'supervised' in §4.2.","section":"Abstract; Figure 1; §4.2"},{"comment":"The table formatting is difficult to parse: check marks are used as row labels, and entries such as '31.78✓ SSIM 0.843' mix values and labels in a single cell. Please reformat so each row is clearly labeled.","section":"Table 6"},{"comment":"The terminology is inconsistent: the abstract describes MSFT as 'one-stage', while the contributions state a 'two-stage integrated network', and §3.2 describes a 'three-scale U-shaped architecture'. Please clarify the stage/scale terminology.","section":"Abstract; §3; §3.2"},{"comment":"The attention formula is nonstandard and ambiguous. As written, the softmax argument is K_i^T Q_i (dimension d_k × d_k) rather than the usual Q_i K_i^T (dimension HW × HW), and the modulated value (bY_i ⊙ V_i) appears before the softmax. Please clarify the intended computation and verify the shapes.","section":"Eq. (9)"}],"recommendation":"reject","confidential_remarks":"The manuscript contains multiple unexplained numerical duplications and internal contradictions in the central experimental tables. Even if these were corrected, the absence of code, weights, and baseline reproduction protocols would prevent independent verification. I do not see a path to acceptance without a full re-evaluation of all baselines and repeated runs under a common protocol."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Quick take: the architecture is a real but incremental recombination of known Fourier/attention blocks, and the headline 11.92 dB gain on SDSD-outdoor is not supported by the paper's own tables.\n\nWhat's actually new: FTG-MSA, which fuses the amplitude spectrum of a max-channel illumination prior with the input feature's amplitude, is a reasonably clean idea that extends FECNet/UHDformer-style Fourier guidance. The MSSA stack (SCSA + MSA + CMUNeXt) is a standard plug-and-play combination. The authors also run 12 independent runs and report stability statistics, which is more than many LLIE papers do, and their limitations section is honest.\n\nWhere it falls down: the empirical evidence is internally inconsistent. Retinexformer's LOL-v1 PSNR is 25.16 in Table 4 but 22.43 in Table 5—exactly Restormer's number. The 'W/O MSA' row in Table 7 is identical to the 'W/O DRA' row in Table 8 on PSNR and nearly identical on SSIM, which is impossible if DRA is a submodule of MSA. The final MSFT scores in Table 4 equal the 12-run means in Table 3 to two decimals on all seven datasets, although the text says the run with highest validation PSNR was selected; that would be a remarkable coincidence unless the mean was reported by mistake. On top of that, SNR-Net and Retinexformer are labeled unsupervised, 3DLUT is labeled semi-supervised, and no code or baseline reproduction protocol is given. None of this looks like misconduct—it looks like careless table-building—but it means the claimed SOTA margins are unverifiable as submitted.\n\nWho should read it: someone working on Fourier-domain LLIE might mine it for the FTG-MSA idea, but shouldn't trust the numbers. My recommendation: a serious editor should send it to peer review anyway, because the architecture is plausible and the problems are concrete and fixable. The reviewer's job is to demand corrected tables, rerun baselines, and code. If the 12-run mean/selection discrepancy is just a reporting error, the method could still be strong. But as it stands, the central claim fails its burden of proof.","headline":"The architecture is a plausible Fourier-attention combination, but the benchmark claims are unravelled by the paper's own table inconsistencies.","tokens_in":23008,"tokens_out":4159,"would_cite":false,"duration_ms":35487,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper introduces MSFT, a supervised U-shaped network for low-light image enhancement that fuses Fourier-domain amplitude information from a Retinex-based illumination estimate into multi-scale self-attention at every stage. It claims th","keywords":["low-light image enhancement","Fourier transform","multi-scale attention","Retinex prior","transformer","amplitude fusion","image restoration"],"falsifier":"Recompute every baseline metric from scratch on the same seven dataset splits with the same crop size, number of iterations, and evaluation metrics, then re-run MSFT; the central claim fails if the 11.92 dB gap on SDSD-outdoor does not persist under a unified protocol.","tokens_in":21866,"feed_emoji":"🌙","tokens_out":4597,"duration_ms":43469,"temperature":0.7,"pith_summary":"The paper introduces MSFT, a supervised U-shaped network for low-light image enhancement that fuses Fourier-domain amplitude information from a Retinex-based illumination estimate into multi-scale self-attention at every stage. It claims this design achieves state-of-the-art results on all seven low-light benchmarks tested, most strikingly on SDSD-outdoor where it reports 41.76 dB PSNR and 0.988 SSIM, surpassing the Retinexformer baseline by 11.92 dB PSNR. The architecture is compact at 1.25 million parameters and one-stage, with ablations attributing most of the gain to the Fourier amplitude guidance and the multi-shape attention module. If the reported comparisons hold, the method offers a practical recipe for simultaneously restoring brightness and texture in difficult low-light footage.","feed_headline":"Fourier attention lifts low-light PSNR to 41.76 dB","feed_subtitle":"A 1.25M-parameter network fuses Retinex brightness priors into transformers, topping seven benchmarks if comparisons hold.","key_machinery":"The central mechanism is the Fourier Transform-Guided Multi-Scale Attention (FTG-MSA) module: it runs an FFT on the low-light feature and on a four-channel Retinex-derived brightness map, sums their amplitudes, applies an inverse FFT, and uses the result to modulate value vectors in self-attention via a Hadamard product. This is supported by the Multi-Shape Synergistic Attention (MSSA) block, which combines SCSA, Dilated Square Attention, Dilated Rectangle Attention, and CMUNeXt at the highest channel scale to re-weight sparse high-dimensional features. The Retinex max-channel illumination map serves as the fourth input channel, providing the brightness prior that the frequency fusion carrie","core_discovery":"On its own terms, the paper's discovery is that brightness information can be transported across the network in the frequency domain. Instead of only processing the low-light RGB image, MSFT computes a Retinex illumination prior from the max channel, concatenates it as a fourth channel, applies a Fourier transform to both the feature map and this prior, and replaces the feature map's amplitude with a weighted sum that carries the prior's brightness while preserving the original phase for structure. This Fourier-guided self-attention is repeated at three scales in a U-Net, and a Multi-Shape Synergistic Attention block further re-weights high-dimensional feature channels. The paper reports tha","pith_inferences":["The same amplitude-fusion trick might generalize to other domains with global illumination priors, such as underwater or nighttime dehazing, though the paper does not test this.","Because the paper notes the model fails on strong-light pollution and requires paired clean data, a natural extension is to replace the paired Retinex prior with a self-supervised or zero-shot estimate and measure how much of the gain survives.","A simpler test than building the full architecture: compute the amplitude-swapped Fourier reconstruction from Figure 1 and check whether the perceptual gain matches what the network achieves; if a closed-form amplitude swap already explains most of the improvement, the attention module may be secondary.","The internal inconsistency in baseline tables suggests a reproducibility check: if Retinexformer's true LOL-v1 score is 22.43 rather than 25.16, MSFT's lead shrinks from 2.32 dB to over 5 dB, changing the headline comparison."],"forward_implications":["If the results hold, MSFT provides a parameter-light (1.25M) supervised alternative to diffusion-based low-light enhancers, suitable for paired-data settings.","The amplitude-fusion attention mechanism could be adopted as a drop-in module in other U-shaped restoration networks for tasks where global brightness priors matter.","The reported SSIM of 0.988 on SDSD-outdoor suggests strong structural fidelity, not just brightness recovery, on dynamic video frames.","Ablations attribute most of the SDSD-outdoor gain (6.78 dB) to the Fourier module, implying frequency-domain guidance is the primary performance driver rather than the transformer backbone alone.","The paper's statistical analysis shows high stability on SDSD (std ≤ 0.03 dB) but weaker stability on SMID and LOL-v2-real (std ≥ 1.38 dB), indicating margins may vary across runs."],"fun_headline_variants":["Fourier-domain attention restores texture in dark images","Multi-scale Fourier fusion boosts low-light PSNR by 11.92 dB","Brightness prior in frequency domain lifts low-light detail","U-Net with Fourier-guided attention tops seven benchmarks","Fourier transform plus attention brightens low-light images"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The comparative results assume that the published numbers for existing methods were produced under the same training and evaluation protocol as MSFT; the manuscript's own tables disagree on at least one baseline (Retinexformer is given a PSNR of 25.16 dB on LOL-v1 in one table and 22.43 dB in another), so if the baselines were not run comparably, the claimed margins could collapse.","fun_headline_variants_meta":{"raw":{"variants":["Fourier-domain attention restores texture in dark images","Multi-scale Fourier fusion boosts low-light PSNR by 11.92 dB","Brightness prior in frequency domain lifts low-light detail","U-Net with Fourier-guided attention tops seven benchmarks","Fourier transform plus attention brightens low-light images"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000271,"raw_usage":{"total_tokens":1500,"prompt_tokens":811,"completion_tokens":689,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":555,"completion_tokens_details":{"reasoning_tokens":605}},"tokens_in":555,"tokens_out":689,"duration_ms":7696,"temperature":1.0,"reasoning_tokens":605,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-07-31T23:18:11.918014+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Recompute every baseline metric from scratch on the same seven dataset splits with the same crop size, number of iterations, and evaluation metrics, then re-run MSFT; the central claim fails if the 11.92 dB gap on SDSD-outdoor does not persist under a unified protocol.","supporting_citations":[],"review_version":1}