{"id":"71a2c15a-2f49-4226-ba04-d135f60a91fa","arxiv_id":"2603.02692","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":4.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":7,"one_line_summary":"A one-step diffusion super-resolution model combining detail-aware loss weighting, latent residual refinement, and tunable low/high-frequency injection reports the best fidelity-detail balance among compared diffusion baselines.","lead":"This paper presents a one-step diffusion model for real-world image super-resolution that adds three ingredients: detail-weighted training, a latent residual refinement block, and tunable frequency injection at inference time. It reports better perceptual and fidelity metrics than existing one-step and multi-step diffusion baselines in a single inference step, at the cost of small margins and a test-set-tuned inference knob.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"LFIM parameters tuned on RealSR test set; Table 1's RealSR row and the claimed SOTA balance are optimistic without re-validation.","rationale":"The reader's weakest_assumption identified LFIM test-set tuning as the key threat to the headline numbers; I agree and believe it is the most load-bearing concern. The central claim of superior real-world SR performance is directly supported by Table 1, and the RealSR row is obtained with hyperparameters selected on that same test set. This is a protocol flaw that can inflate perceptual-fidelity metrics and the gap to PiSA-SR. It also colors DRealSR/DIV2K rows, which reuse the test-selected configuration. The reader's other concerns (LPIPS used as both training objective and evaluation metric, no code/checkpoints, no variance statistics) are valid and compound the issue. I additionally noticed a concrete internal inconsistency in Table 3: the RealSR row reports LRRB MSE 0.1045 > baseline 0.1029, yet labels it a 1.62% improvement, and the column averages do not match the listed values. This is not the primary load-bearing issue for the headline result, but it reinforces the need for code release and corrected tables. A CONDITIONAL verdict remains apt: authors should release code/checkpoints, re-run LFIM tuning on a validation split, report variance, and fix Table 3. No change to the reader's verdict is needed.","tokens_in":19491,"tokens_out":14841,"duration_ms":139519,"concrete_test":"Select lf_alpha and hf_beta using only a validation split of training data (or a held-out half of RealSR), then re-evaluate FiDeSR on the untouched RealSR test half. If the chosen parameters are not (0.2,0.2) or the re-evaluated metrics fall short of Table 1's RealSR row (LPIPS 0.2626, MANIQA 0.6681, FID 109.68), the reported gains are partially due to test-set tuning.","verdict_should_be":"UNCHANGED","load_bearing_attack":"Table 1's RealSR row—and, indirectly, the DRealSR/DIV2K rows—are produced with LFIM strengths (lf_alpha=0.2, hf_beta=0.2) that were selected by scanning injection intensities on the RealSR test set itself (Supp. Sec. G; Table 8). The same RealSR test set then appears as a headline row in Table 1. This makes the RealSR comparison a test-set-tuned result: the reported PSNR 26.02, LPIPS 0.2626, FID 109.68, etc., are the best (or a favored) working point from a scan on the evaluation set, not a prediction of the trained model at a pre-specified setting. The DRealSR and DIV2K rows use the same configuration, so they inherit a choice made with knowledge of a same-domain test set. If this protocol fails, the gap between FiDeSR and PiSA-SR (e.g., RealSR LPIPS 0.2626 vs 0.2672; FID 109.68 vs 124.18) may be partly an artifact of test-set tuning rather than a property of the one-step model. The paper's central claim of superior real-world SR performance therefore rests on an insecure empirical basis. A secondary quantitative inconsistency appears in Table 3 (RealSR LRRB MSE 0.1045 > baseline 0.1029 yet listed as a 1.62% improvement; averages also do not match), which further underscores that the reported numbers need independent verification.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper proposes FiDeSR, a one-step latent diffusion super-resolution framework that combines three components: a Detail-aware Weighting (DAW) strategy that reweights training losses by a spatial difficulty map, a Latent Residual Refinement Block (LRRB) that refines the predicted residual, and a Latent Frequency Injection Module (LFIM) that injects low- and high-frequency components at inference. The authors claim that FiDeSR achieves a better perceptual-fidelity balance than existing one-step and even multi-step diffusion SR methods, with strong results on RealSR, DRealSR, and a DIV2K-based synthetic test set, while requiring only one diffusion step.","tokens_in":19820,"tokens_out":4838,"duration_ms":46428,"significance":"If the empirical claims survive independent validation, this is a useful contribution to efficient real-world SR: it suggests that a one-step latent diffusion model can improve both perceptual quality and structural fidelity through loss reweighting, residual refinement, and tunable frequency injection. The ablations in Table 4 show directionally sensible and monotonic effects of LF/HF injection, and the LRRB/DAW ablations in Table 2 indicate that the training-side components help on no-reference metrics. The paper is generally well structured and the proposed modules are clearly described. However, the headline comparisons rest on an evaluation-protocol concern (inference-time parameters selected on the test set) and on a quantitative inconsistency in Table 3 that need to be resolved before the central claim can be accepted.","major_comments":[{"comment":"The LFIM injection strengths lf_alpha=0.2 and hf_beta=0.2 were selected by scanning on the RealSR test set, and the same RealSR test set is then reported as the headline row in Table 1. This makes the RealSR comparison a test-set-tuned result, and the DRealSR and DIV2K rows inherit hyperparameters chosen with knowledge of a same-domain test set. The reported advantages over PiSA-SR on RealSR (e.g., FID 109.68 vs 124.18) may therefore be partly an artifact of tuning rather than a property of the one-step model. Please re-select the injection strengths on a held-out validation split (or pre-specify them before evaluation) and re-run all benchmark rows; if the parameters are meant to be user-adjustable controls, the state-of-the-art claim should be qualified accordingly.","section":"Sec. 4.2 (Table 1) and Supp. Sec. G (Table 8)"},{"comment":"The RealSR row is internally inconsistent: the baseline MSE is 0.1029 and the LRRB MSE is 0.1045, i.e., LRRB increases the error, yet the table reports a 1.62% improvement. The average row also does not match the dataset rows (0.1049 vs 0.1032). This contradicts the claim in Sec. 4.3 that the LRRB-equipped model consistently reduces the high-frequency noise prediction error. Because Table 3 is the quantitative evidence for the LRRB component, this inconsistency is load-bearing and must be corrected and independently verified.","section":"Sec. 4.3, Table 3"},{"comment":"DAW constructs its perceptual error map using LPIPS, and the training loss includes a spatially weighted LPIPS term; LPIPS is also a headline evaluation metric. The paper should clarify the degree to which the reported LPIPS gains reflect direct per-pixel optimization rather than generalizable perceptual improvement. Currently Table 2 reports only no-reference metrics for the DAW ablation and therefore does not disentangle this. Please add an ablation with p=0 (or unweighted LPIPS) and report the full metric set, including LPIPS, DISTS, and FID, to demonstrate that the gains are not solely due to optimizing the evaluation metric.","section":"Eqs. (5), (6), (10) and Table 2"}],"minor_comments":[{"comment":"The training set includes DIV2K and the synthetic test set is cropped from DIV2K-validation. Please state explicitly that the training split excludes the validation images used for testing, to rule out train/test overlap.","section":"Sec. 4.1"},{"comment":"The coefficients p, w_max, and alpha appear in the DAW implementation but are not given values in the main text. Report the chosen values and, if possible, a sensitivity study.","section":"Alg. 1 and Sec. 3.3"},{"comment":"The detail map D is defined as the mean response of Sobel, Laplacian, and variance filters, but the three operators produce outputs on very different scales. Clarify whether each operator is normalized before averaging.","section":"Eq. (3)"},{"comment":"The caption and table use 'PISA-SR' while the rest of the paper uses 'PiSA-SR'. Please standardize the name.","section":"Table 6"},{"comment":"The notation Δ_LP and Δ_HP is introduced without formal definition. Define these quantities explicitly and specify how the Butterworth filters and the spatial/channel gates are computed.","section":"Sec. 3.6"}],"recommendation":"major_revision","confidential_remarks":"The main risk is the LFIM hyperparameter selection on the RealSR test set; if the authors cannot provide an independent validation-based selection, the headline SOTA claims on real-world benchmarks are substantially weakened. The Table 3 inconsistency also needs a careful check. I would be willing to re-review a revised version that addresses these two issues."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Short version: this is a competent engineering paper that makes a plausible one-step SR recipe, but the headline comparison is partly a test-set-tuned result, so the main claim of a new fidelity-detail Pareto point is not yet established.\n\nWhat is genuinely new here is the specific combination: replacing the PiSA-SR global residual with a latent refinement block, adding a detail-aware weighting during training, and then injecting low/high-frequency components at inference. The three components each have clear precedents, but the composition is not in the literature, and the ablations are better than most SR papers. Table 2 shows LRRB and DAW each help, and importantly those numbers are evaluated before LFIM, so they are not contaminated by the inference-time tuning. Table 4's monotonic behavior — LF injection improving PSNR/SSIM, HF injection improving MUSIQ/MANIQA — is exactly what you would expect if the mechanism works as described. That internal consistency earns real credit.\n\nThe soft spot is the one the stress test flags. The LFIM strengths (lf_alpha=0.2, hf_beta=0.2) were chosen by scanning on the RealSR test set, and the same RealSR test set appears as a headline row in Table 1. That makes the RealSR row a tuned result, not a prediction, and the DRealSR and DIV2K rows inherit a configuration selected with knowledge of a same-domain test set. So the gap to PiSA-SR — especially the FID difference of 124.18 to 109.68 — may partly be an artifact of protocol. This is the load-bearing flaw for the paper's central claim.\n\nThere is also a minor but real numerical inconsistency in Table 3: RealSR MSE moves from 0.1029 to 0.1045, yet the table lists that as a 1.62% improvement, and the average row does not match. Could be a typo, but it makes me want to re-verify every number. And the fact that LPIPS is both a training objective and a headline reported metric means part of the perceptual gain is by construction.\n\nNo code, no checkpoints, and no variance or significance statistics are provided, so the magnitude of the reported advantages cannot currently be checked.\n\nWho is this for? Someone working on one-step diffusion SR who wants the component-level recipe, and anyone teaching a methods class on why inference-time tuning on the test set invalidates a SOTA claim. I would send it to peer review, but with a clear request to fix the protocol: either pick the injection strengths on a validation split, or report the full sweep as the result and refrain from cherry-picking the balanced point against other methods. The LRRB and DAW contributions may survive that correction; the current LFIM-based claims will not.","headline":"Useful one-step SR engineering, but the headline numbers are partly tuned on the test set — the central SOTA claim needs re-validation before I'd trust it.","tokens_in":110,"tokens_out":1429,"would_cite":false,"duration_ms":24004,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"One-step diffusion super-resolution can raise perceptual quality and structural fidelity together, when training is error-weighted and the latent is frequency-refined at inference.","keywords":["image super-resolution","one-step diffusion","latent diffusion models","detail-aware weighting","residual refinement","frequency injection","real-world super-resolution","perception-distortion trade-off"],"falsifier":"Re-evaluate FiDeSR on DRealSR and RealSR with LFIM parameters frozen at zero (or chosen only on a held-out validation split) and compare against PiSA-SR and OSEDiff on LPIPS, DISTS, and FID; if the margins shrink to noise or reverse — for instance if DRealSR LPIPS climbs back above 0.30 or FID above 135 — the reported balance of perceptual quality and fidelity is an artifact of test-set tuning rather than of the trained network.","tokens_in":19312,"feed_emoji":"⚡","tokens_out":10494,"duration_ms":86413,"temperature":0.7,"pith_summary":"The paper sets out to establish that the perception–distortion trade-off that has forced diffusion-based super-resolution to run tens or hundreds of denoising steps can be broken with a single step: FiDeSR's restored images are reported to be both more faithful in structure and more realistic in texture than those of existing diffusion methods. It diagnoses why one-step diffusion SR fails — low-frequency inconsistency from latent compression and under-produced high-frequency detail from truncated denoising — and addresses each failure with a named mechanism. Detail-aware Weighting concentrates training loss on regions where prediction error is largest; the Latent Residual Refinement Block corrects the diffusion network's coarse residual prediction; and the Latent Frequency Injection Module reinjects low- and high-frequency content at inference with user-adjustable strength. On real-world benchmarks the paper reports the best perceptual scores among nine diffusion methods at a single diffusion step, suggesting that efficient one-step restoration need not surrender fidelity.","feed_headline":"One diffusion step beats 200-step rivals on perceptual quality","feed_subtitle":"FiDeSR pairs error-weighted training with latent frequency injection to hold structure while adding detail.","key_machinery":"The load-bearing object is the residual latent restoration identity zr = zL − r', which reduces super-resolution to predicting how much to subtract from the low-quality latent. FiDeSR upgrades this identity in two ways. LRRB replaces the single coarse residual with a refined one, r' = r + ∆r, where ∆r is produced by a stack of residual-in-residual dense blocks operating on the concatenation of the LQ latent and the initial residual — a learned correction that directly targets the instability of one-step noise prediction. LFIM then modifies the latent itself at inference: z ← z + α·M_sp·M_ch·∆_LP for low frequencies, and z ← z + β·M_sp·M_ch·∆_HP for high frequencies, where M_sp and M_ch are s","core_discovery":"The central claim, stated on the paper's own terms, is that a one-step latent diffusion model can restore images with both high perceptual quality and faithful content — the combination prior one-step methods sacrificed. The paper locates the failure in the residual latent formulation z0 = zL − r: a single globally predicted residual is unstable and leaves high-frequency noise uncorrected, while the latent encoder discards low-frequency structure. FiDeSR's answer is three mechanisms acting at three stages. During training, Detail-aware Weighting (DAW) builds a per-pixel difficulty map W = D ⊙ E from a spatial detail map (Sobel, Laplacian, local variance of the ground truth) times an error ma","pith_inferences":["Editorial: the LFIM working point (lf_alpha=0.2, hf_beta=0.2) was selected by scanning injection strengths on the RealSR test set, and that same set produces the headline RealSR row in Table 1; a skeptic should read those numbers as the ceiling of a tuned module, not of the trained network alone.","Editorial: if the diagnosis is right — latent compression causes low-frequency drift and one-step truncation causes high-frequency loss — then LFIM-style frequency reinjection could serve as a drop-in inference module for other latent-diffusion restoration tasks, such as deblurring or face restoration, regardless of training objective.","Editorial: a natural testable extension is to replace the global constants α and β with per-image or per-patch selection, e.g., optimizing them against a no-reference quality score; the reported monotonicity of PSNR/SSIM in α and of MUSIQ/MANIQA in β suggests such an autotuner would stay on a smooth surface.","Editorial: because the DAW error map uses an LPIPS backbone, the reported fidelity gains may inherit that metric's bias; retraining with a different perceptual error map is a direct robustness check the paper does not run."],"forward_implications":["If FiDeSR holds up, real-world super-resolution no longer needs multi-step denoising for fidelity: its single-step output beats 20–200-step methods on perceptual metrics while running roughly 25–100 times faster.","The LFIM injection strengths (lf_alpha, hf_beta) provide a no-retraining control dial between distortion-oriented and perception-oriented outputs, so one model can serve both archival-fidelity and display-sharpness use cases.","The measured reduction in high-frequency noise prediction error (1.24–1.99% across datasets) indicates LRRB addresses a general weakness of one-step distillation rather than a data-specific artifact, and could be reused in other single-step generative restoration pipelines.","The FID gains and user-study votes imply restored images sit closer to the true image distribution as judged both by statistics and by human preference."],"fun_headline_variants":["One-step diffusion SR that keeps details faithful","Residual-refined one-step diffusion for high-fi SR","FiDeSR: one-step diffusion with detail-preserving fidelity","Error-weighted training powers faithful one-step diffusion SR","Detail-aware one-step diffusion for faithful super-resolution"],"cache_read_input_tokens":2304,"weakest_assumption_plain":"The headline numbers assume it is legitimate to tune the inference-time injection strengths (lf_alpha, hf_beta) on the RealSR evaluation set and then report that same set's improved scores; if the working point must be fixed without seeing test data, the claimed advantage over the next-best one-step method is optimistic.","fun_headline_variants_meta":{"raw":{"variants":["One-step diffusion SR that keeps details faithful","Residual-refined one-step diffusion for high-fi SR","FiDeSR: one-step diffusion with detail-preserving fidelity","Error-weighted training powers faithful one-step diffusion SR","Detail-aware one-step diffusion for faithful super-resolution"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000791,"raw_usage":{"total_tokens":3304,"prompt_tokens":705,"completion_tokens":2599,"prompt_tokens_details":{"cached_tokens":256},"prompt_cache_hit_tokens":256,"prompt_cache_miss_tokens":449,"completion_tokens_details":{"reasoning_tokens":2524}},"tokens_in":449,"tokens_out":2599,"duration_ms":17601,"temperature":1.0,"reasoning_tokens":2524,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-02T19:18:12.774776+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-evaluate FiDeSR on DRealSR and RealSR with LFIM parameters frozen at zero (or chosen only on a held-out validation split) and compare against PiSA-SR and OSEDiff on LPIPS, DISTS, and FID; if the margins shrink to noise or reverse — for instance if DRealSR LPIPS climbs back above 0.30 or FID above 135 — the reported balance of perceptual quality and fidelity is an artifact of test-set tuning rather than of the trained network.","supporting_citations":[],"review_version":1}