{"id":"b9af0fb9-9eba-46ca-b1c4-b28c4c37d5d4","arxiv_id":"2504.18524","paper_version":1,"verdict":"CONDITIONAL","confidence":"HIGH","novelty_score":6.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"NR-IQA-guided sampling and LoRA-regularized fine-tuning improve reported NR-IQA scores over the human-feedback HGGT baseline without using human labels.","lead":"The paper tests whether automated no-reference image quality scores can replace human raters when training super-resolution models. It reports that quality-score-based sampling plus fine-tuning beats a human-annotated baseline on several quality metrics, while the largest gains appear on the very metric being optimized.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"MUSIQ is both the training signal and the headline evaluation metric, and the only human check is a 12-rater, single-architecture study; the claim that automatic IQA matches human feedback is not yet disentangled from metric overfitting.","rationale":"The reader's weakest assumption identifies MUSIQ's validity as a human-quality proxy, which is indeed the load-bearing premise. My stress-test sharpens this into a specific overfitting risk: MUSIQ is simultaneously the selection signal, the optimization target, and the primary evaluation metric, and the AMO+FT result surpasses the paper's own soft gold-standard bound on MUSIQ, exactly the signature of metric overfitting rather than evidence of genuine quality gain. The paper does include independent support: a modest user study, complementary NR-IQA metrics, and an honest limitations section. However, the user study is small and covers only one architecture, and the complementary NR-IQA metrics are themselves learned proxies that correlate with human opinion but are not human judgment. The RealESRGAN results, where AMO+FT does not surpass UPos+FT, further narrow the scope of the claim. None of this invalidates the paper's empirical contribution, but it means the central claim should remain conditional pending a more decisive human evaluation. I therefore recommend no change to the reader's CONDITIONAL verdict.","tokens_in":25095,"tokens_out":6308,"duration_ms":64937,"concrete_test":"Pre-register a larger human study (at least 50 naive raters, covering both SwinIR and RealESRGAN on HGGT Test-100) that obtains per-image paired preferences for AMO+FT versus UPos and for AMO+FT versus the highest-MUSIQ human-positive GT. Then compute the correlation between per-image MUSIQ gain (AMO+FT minus UPos) and per-image human preference. If the correlation is weak or negative, or if raters prefer the positive GT over AMO+FT, the MUSIQ proxy assumption fails and the central claim is unsupported.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim requires that maximizing MUSIQ improves true perceptual quality, not just the metric. That premise is load-bearing because MUSIQ enters twice: as the AMO selection signal in §4.2 and as the optimization objective in Eq. 5 of §4.3, while Table 5 lists MUSIQ as the first NR-IQA evaluation metric. The paper's own Limitations section (§14) concedes that higher IQA scores do not guarantee improved human perceptual quality. Three concrete observations sharpen the concern. (1) SwinIR-AMO+FT reaches MUSIQ 70.81, above the paper's own 'Gold Standard' of 69.64 computed from the best human-derived GT per quintuplet; since the model is trained to maximize this exact score, surpassing the GT bound is precisely what Goodharting predicts. (2) Naive MUSIQ optimization produces adversarial structured artifacts (§4.3, Fig. 3, Supp. §10.2); LoRA reduces but does not eliminate them, as Fig. 8 still shows line-orientation and aliasing artifacts. (3) The only human check is a 12-rater, single-architecture (SwinIR) preference test against UPos; there is no human validation for RealESRGAN, no comparison against the positive GTs or gold-standard images, and no check of whether preference tracks MUSIQ margins. The possibility that the method mainly games MUSIQ while shifting the sharpness/artifact tradeoff has therefore not been excluded. The stated claim that automatic IQA yields results 'perceptually on par or better' than human-feedback fine-tuning is not yet established at the confidence implied.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper studies whether no-reference image quality assessment (NR-IQA) models can replace human feedback in super-resolution training. It first evaluates a large set of NR-IQA metrics on two human-preference SR datasets (SBS180K and HGGT), identifying MUSIQ as the most aligned metric. It then proposes two mechanisms: (i) reweighted sampling of multiple ground-truth targets based on IQA scores (SMA, SMP, and AMO in §4.2), and (ii) fine-tuning with a differentiable NR-IQA loss regularized by LoRA (§4.3, Eq. 5). The combination AMO+FT is reported to outperform the human-feedback-based HGGT 'UPos' baseline on NR-IQA metrics and in a 12-participant user study (Supp. §12), supporting the claim that automatic IQA can match or exceed human-guided fine-tuning for SR.","tokens_in":25438,"tokens_out":5100,"duration_ms":49351,"significance":"If established, the result would be significant: it would show that modern NR-IQA metrics can act as a scalable automatic reward for perceptual SR optimization, avoiding costly human annotation. The paper is also valuable for its systematic comparison of 42 NR-IQA variants on human SR preference data, for the additional RealSR generalization results, and for the honest ablation showing that direct IQA optimization naively produces adversarial artifacts (Fig. 3) while LoRA mitigates them. However, the central 'perceptually on par or better than human feedback' claim is not yet airtight, because MUSIQ is simultaneously the optimization target and the headline evaluation metric, and the only human check covers one architecture with 12 raters.","major_comments":[{"comment":"The largest reported gains are on MUSIQ itself, which is both the training signal and the primary evaluation metric: for SwinIR-AMO+FT vs SwinIR-UPos, MUSIQ increases by 4.42 points while NIMA, Q-Align, and TOPIQ increase by 0.13, 0.19, and 0.08, respectively. Those non-optimized metrics are the only direct evidence of perceptual generalization, and their small magnitude is not discussed with significance or confidence intervals. Please report variances or error bars across test images, and add an evaluation metric that was never used during training and is not a monotone transform of MUSIQ, to break the circularity.","section":"Table 5, Eq. (5), §4.2"},{"comment":"SwinIR-AMO+FT reaches MUSIQ 70.81, above the 'Gold Standard' value of 69.64 computed by taking the best MUSIQ score among the original and enhanced GTs per quintuplet. Since the fine-tuning loss maximizes MUSIQ directly, surpassing the GT bound is exactly what metric overfitting (Goodharting) would predict; the manuscript currently presents this as evidence of success without addressing the alternative explanation. Please add a human evaluation or an analysis of the learned artifacts that supports the interpretation that the model truly produces higher-quality images than the best available GT.","section":"§5.1, Table 5, Gold Standard row"},{"comment":"The user study is too narrow to carry the general claim: it uses only 12 raters, only the SwinIR architecture, only the AMO+FT vs UPos comparison, and no comparison against the gold-standard GTs or the original HR images. In particular, there is no human validation for RealESRGAN, no test of whether preference is driven by the same artifacts flagged in Supp. §14 and Fig. 8, and no analysis of whether the 69.7% preference correlates with the MUSIQ margin. A larger, architecture-diverse human study with artifact-focused instructions is needed before claiming 'perceptually on par or better than SISR finetuned with human feedback'.","section":"Supp. §12"},{"comment":"The paper's own limitations concede that higher IQA scores do not guarantee improved human perceptual quality and show aliasing-like line artifacts and mangled text under AMO+FT. Combined with the known susceptibility of NR-IQA models to adversarial perturbations (cited in §4.3), this means the possibility that AMO+FT partially games MUSIQ while shifting the sharpness/artifact tradeoff has not been excluded. Please add an artifact-level analysis (e.g., side-by-side with original GT, text-specific SR metrics, or a targeted human study) to rule this out.","section":"Supp. §14, Fig. 8"}],"minor_comments":[{"comment":"Clarify the temperature convention in SoftMax_τ: with the standard softmax, τ→0 yields argmax and τ→∞ uniform, matching the text, but the notation is easy to misread; please define the formula explicitly, e.g., SoftMax_τ(Q) = exp(Q/τ)/Σ exp(Q/τ).","section":"§4.2"},{"comment":"In the RealESRGAN paragraph, fix the typo 'neglible' and the inconsistent expansion 'DIST' vs 'DISTS'.","section":"§5.1"},{"comment":"Add the ranges and directions of the NR-IQA metrics to the main-table caption (they are currently only in Supp. Table 7), so that absolute values like MUSIQ 70.81 vs 69.64 are interpretable.","section":"Table 5 caption"},{"comment":"The statement that negatives are rare (∼6%) and MUSIQ ranks a negative highest in ∼34% of such cases is useful; consider moving it to the main text, since it is part of the justification for using MUSIQ despite poor NM in Table 4.","section":"Supp. §7.1"},{"comment":"Please provide any rationale or sensitivity analysis for the LoRA rank choices (rank 48 for SwinIR, 24 for RealESRGAN), since these could affect the artifact behavior observed in Fig. 8.","section":"Supp. §8.3"},{"comment":"State explicitly that the negative sign before λ_Q Q(I) reflects that Q is higher-is-better, so the loss is minimized by increasing Q.","section":"Eq. (5)"}],"recommendation":"major_revision","confidential_remarks":"The central concern is that the paper's headline claim rests in part on a circular evaluation, but the authors have already included multiple non-optimized metrics and a user study, so the issue is fixable within a revision. I recommend major revision rather than rejection. The paper is within scope and likely to be of interest to the SR and IQA communities; the main risk is overclaiming the perceptual equivalence to human-feedback fine-tuning."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"I read this one carefully, and it is better than the typical SR training trick paper. The genuinely new parts are the metric analysis and the way they use the results. Running 42 NR-IQA variants on SBS180K and HGGT, then reporting accuracy, complementarity, and misalignment rates, is real work that will be useful to anyone in the field. The choice of MUSIQ is justified rather than assumed, and the observation that NIMA and Q-Align complement MUSIQ is a nice practical finding. The method itself is simple and clearly described: reweighted multi-GT sampling (SMA, SMP, AMO) plus LoRA-regularized fine-tuning on a differentiable quality score. AMO+FT beats the human-annotation baseline UPos on all four reported NR-IQA metrics, not just MUSIQ, and wins a 12-rater user study by 69.7% with p<0.01. That is real evidence. The RealSR results and the ablation with PaQ-2-PiQ also help. And the limitations section is honest: it explicitly says higher IQA scores do not guarantee better human perceptual quality and shows examples of line-orientation artifacts and mangled text. The soft spot is the one the stress test highlights. MUSIQ is doing triple duty: it selects the patches in AMO, it is the optimization objective in Eq. 5, and it is the first metric in Table 5. The biggest gain is on MUSIQ itself (SwinIR AMO+FT is 70.81, above the paper's own gold-standard value of 69.64), while the gains on NIMA, Q-Align, and TOPIQ are small. That pattern is exactly what you would expect from partial Goodharting. They try to defend against it with complementary metrics and the user study, but the user study is 12 raters and only covers SwinIR. So the strong claim that automatic IQA is on par with human feedback is plausible, not yet nailed down. I would not call it a load-bearing flaw, because the paper is transparent about it and the non-optimized metrics do move in the right direction, but the title and abstract overstate the confidence. Also: no code released, and the baseline set is narrow. They compare against HGGT UPos but not against a modern diffusion-based SR model or a strong GAN baseline trained without HGGT. That limits how far the 'state of the art' claim can go. Who is this for? People working on perceptual SR or on using IQA as a training signal. It will get them up to speed on which NR-IQA metrics align with human judgment for SR and give them a sensible recipe. It deserves a serious referee. I would send it out, and the review should push for a larger user study, an evaluation where a different NR-IQA model is optimized to test generalization, and code release. With those, this could be a solid conference paper.","headline":"A worthwhile empirical study of using NR-IQA to guide SR training, with an honest limitations section; the main claim is plausible but the circularity between the optimized metric and the headline evaluation metric means it should be read with that caveat.","tokens_in":826,"tokens_out":883,"would_cite":true,"duration_ms":24829,"reading_group":"yes","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"An automatic no-reference image-quality measure can replace human raters in super-resolution fine-tuning, matching or beating human-feedback training without any manual labels.","keywords":["single-image super-resolution","no-reference image quality assessment","perceptual quality","human feedback","MUSIQ","LoRA fine-tuning","perception-distortion tradeoff","multi-ground-truth supervision"],"falsifier":"A content-diverse user study in which human raters consistently prefer the human-feedback baseline over AMO+FT outputs, or prefer images with lower MUSIQ scores, would refute the central claim; a sharper test is to search for SR image pairs with similar MUSIQ scores where one is visibly artifact-laden, since such pairs would show the score is not a dependable quality proxy.","tokens_in":24890,"feed_emoji":"🖼️","tokens_out":11023,"duration_ms":91709,"temperature":0.7,"pith_summary":"The paper argues that a pretrained no-reference image-quality assessment (NR-IQA) model can stand in for human preference feedback when training single-image super-resolution (SISR) models. On two human-labeled SR datasets, the authors test which NR-IQA metrics align best with human judgments and select MUSIQ as the most reliable. They then use MUSIQ in two ways: to reweight which of several candidate high-resolution targets the network is trained on, and as a differentiable objective in a LoRA-regularized fine-tuning step. The combined method, called AMO+FT, is evaluated on the held-out HGGT Test-100 benchmark, where it reports higher perceptual-quality scores than the human-feedback baseline on all four NR-IQA metrics tested and was preferred by human raters 69.7% of the time in a 12-rater study. If correct, the practical payoff is that perceptually strong super-resolution can be trained without expensive manual annotation.","feed_headline":"No-reference IQA replaces human raters in super-resolution training","feed_subtitle":"Sampling and LoRA-regularized optimization with MUSIQ beat human-feedback training on all NR metrics and a user study.","key_machinery":"The central object is MUSIQ (Multi-Scale Image Quality Transformer), a learned no-reference image quality model, used twice: as a scoring function for choosing training targets and as a differentiable loss for fine-tuning. The paper first analyses 42 NR-IQA variants to justify selecting MUSIQ for its human alignment and low positive-misalignment rate. In sampling, MUSIQ weights a softmax distribution over candidate ground truths, and in the 'argmax-online' (AMO) variant it selects the highest-quality patch among several super-resolved candidates at training time. In fine-tuning, adding $\\lambda_Q Q(\\hat{I})$ to the loss would normally act as an adversarial attack on the quality network, so the paper restricts updates to low-rank adaptation (LoRA) parameters, which removes the structured artifacts while still letting the model improve its quality score.","core_discovery":"The paper's central claim is that pretrained NR-IQA models can act as a fully automatic substitute for human rankings in perceptual super-resolution, and that combining two uses of them—quality-weighted sampling of multiple ground truths and regularized direct optimization—beats the human-feedback baseline while requiring no human annotations. On the held-out HGGT Test-100 set, the AMO+FT combination raises the MUSIQ score from 66.39 to 70.81 for SwinIR and from 65.93 to 71.67 for RealESRGAN, improves all four NR metrics over the UPos baseline, and wins the preference study with p<0.01, while taking only small penalties on mid-level reference metrics such as LPIPS and DISTS and larger losses on pixel-level PSNR and SSIM. The paper interprets this as moving to a more human-centric point on the perception-distortion tradeoff, prioritizing standalone image quality over pixelwise fidelity.","pith_inferences":["Beyond the paper, the reported NR-IQA gains are partly self-referential because MUSIQ is both the training signal and the lead evaluation metric; the paper's own Limitations section concedes that higher IQA scores do not guarantee improved human perceptual quality, so the complementary metrics and the user study carry extra weight.","Beyond the paper, the same sampling-plus-optimization recipe could transfer to other restoration tasks with subjective quality targets, such as denoising, deblurring, or compression-artifact removal, since no paired ground truth is needed at deployment.","Beyond the paper, a testable extension is to train or fine-tune an NR-IQA model specifically on super-resolution artifacts, or to combine multiple IQA signals, which the paper suggests could reduce the structured artifacts seen under naive optimization."],"forward_implications":["Super-resolution training can proceed without any human ranking of candidate ground truths; the automatic NR-IQA signal is sufficient to match or exceed the human-feedback baseline on perceptual metrics.","The perception-distortion tradeoff shifts: NR-IQA fine-tuning lowers pixel-level PSNR and SSIM while raising high-level perceptual scores, but unlike naive optimization it preserves or improves shift-tolerant mid-level similarity (LPIPS-ST).","The combination of online argmax sampling and LoRA-regularized fine-tuning gives the largest gains, more than either component alone, on both SwinIR and RealESRGAN architectures.","Which NR-IQA model is chosen matters: replacing MUSIQ with PaQ-2-PiQ degrades complementary perceptual metrics, so selecting a human-aligned metric is part of the recipe.","A 12-rater user study prefers the automatic AMO+FT outputs over the human-guided UPos baseline 69.7% of the time, with the rater mean significantly above 50% (p<0.01)."],"supporting_citations":[{"why":"Supplies the HGGT dataset and the UPos human-feedback baseline that the method builds on and compares against.","marker":"[12]"},{"why":"MUSIQ is the NR-IQA model selected for reweighted sampling and for the direct optimization objective.","marker":"[41]"},{"why":"The SBS180K human-preference dataset is used to benchmark which NR-IQA metrics align with human judgments.","marker":"[43]"},{"why":"LoRA low-rank adaptation is the regularizer that prevents naive NR-IQA optimization from producing adversarial artifacts.","marker":"[36]"},{"why":"Supplies the direct fine-tuning-on-differentiable-rewards approach that motivates the optimization objective.","marker":"[16]"},{"why":"Real-ESRGAN is one of the two SR architectures and degradation settings used for training and evaluation.","marker":"[85]"},{"why":"SwinIR is the other SR architecture used for training and evaluation.","marker":"[50]"},{"why":"NIMA is used as a complementary NR evaluation metric and in the metric-alignment analysis.","marker":"[74]"},{"why":"Q-Align is used as a complementary NR evaluation metric and in the metric-alignment analysis.","marker":"[88]"},{"why":"TOPIQ is used as an additional NR evaluation metric and has the second-best positive-misalignment rate after MUSIQ.","marker":"[11]"}],"fun_headline_variants":["NR-IQA models replace human raters for super-resolution","Quality-aware SR training beats human feedback without raters","Pretrained IQA scores guide SR, outdoing human-annotated training","Automatic quality scores train super-resolution better than humans","MUSIQ-driven sampling and tuning win user study in SR"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing premise is that MUSIQ's score reliably tracks human perceptual quality on super-resolved images, a premise the paper's own Limitations section concedes is not guaranteed, so that maximizing it improves perceived quality rather than merely gaming the metric.","fun_headline_variants_meta":{"raw":{"variants":["NR-IQA models replace human raters for super-resolution","Quality-aware SR training beats human feedback without raters","Pretrained IQA scores guide SR, outdoing human-annotated training","Automatic quality scores train super-resolution better than humans","MUSIQ-driven sampling and tuning win user study in SR"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000404,"raw_usage":{"total_tokens":2100,"prompt_tokens":941,"completion_tokens":1159,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":557,"completion_tokens_details":{"reasoning_tokens":1075}},"tokens_in":557,"tokens_out":1159,"duration_ms":10144,"temperature":1.0,"reasoning_tokens":1075,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-16T10:13:29.633618+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"A content-diverse user study in which human raters consistently prefer the human-feedback baseline over AMO+FT outputs, or prefer images with lower MUSIQ scores, would refute the central claim; a sharper test is to search for SR image pairs with similar MUSIQ scores where one is visibly artifact-laden, since such pairs would show the score is not a dependable quality proxy.","supporting_citations":[{"cited_title":"MUSIQ: Multi-scale image qual- ity transformer","cited_arxiv_id":null,"evidence_quote":"MUSIQ is the NR-IQA model selected for reweighted sampling and for the direct optimization objective."},{"cited_title":"Neural side-by-side: Predicting human preferences for no- reference super-resolution evaluation","cited_arxiv_id":null,"evidence_quote":"The SBS180K human-preference dataset is used to benchmark which NR-IQA metrics align with human judgments."},{"cited_title":"LoRA: Low-rank adaptation of large language models","cited_arxiv_id":null,"evidence_quote":"LoRA low-rank adaptation is the regularizer that prevents naive NR-IQA optimization from producing adversarial artifacts."},{"cited_title":"Real-ESRGAN: Training real-world blind super-resolution with pure synthetic data","cited_arxiv_id":null,"evidence_quote":"Real-ESRGAN is one of the two SR architectures and degradation settings used for training and evaluation."},{"cited_title":"SwinIR: Image restoration using swin transformer","cited_arxiv_id":null,"evidence_quote":"SwinIR is the other SR architecture used for training and evaluation."},{"cited_title":"NIMA: Neu- ral image assessment","cited_arxiv_id":null,"evidence_quote":"NIMA is used as a complementary NR evaluation metric and in the metric-alignment analysis."},{"cited_title":"Q- Align: Teaching LMMs for visual scoring via dis- crete text-defined levels","cited_arxiv_id":null,"evidence_quote":"Q-Align is used as a complementary NR evaluation metric and in the metric-alignment analysis."}],"review_version":1}