{"id":"dcc2f983-8f7f-4800-9335-cfd32499905d","arxiv_id":"2605.14847","paper_version":1,"verdict":"UNVERDICTED","confidence":"LOW","novelty_score":7.0,"correctness_risk":"unknown","formal_verification":"none","parameter_count":0,"one_line_summary":"The paper defines perceptual artifact prominence via crowdsourcing, releases the SR-Prominence dataset suite of 3935 masks, and reports that SSIM and DISTS correlate better with human-noticed artifacts than no-reference metrics or specialized detectors.","lead":"This paper introduces artifact prominence, measured as the fraction of viewers who notice a highlighted super-resolution artifact, and releases a crowdsourced dataset of nearly 4000 such annotations across multiple image sources. A smart generalist might read it because current AI image upscalers are judged by metrics that often ignore how noticeable their flaws actually are to people.","discovery_kind":"unclear","skeptic_critique":{"model":"grok-4.3","headline":"Crowdsourced majority-vote prominence labels lack reported stability checks, especially for no-GT Urban100-HR, making metric correlations fragile.","rationale":"The reader’s weakest assumption is exactly the load-bearing point; nothing in the abstract supplies the missing reliability statistics, so the UNVERDICTED verdict stands.","tokens_in":1771,"tokens_out":302,"duration_ms":17091,"concrete_test":"Partition the DeSRA and Urban100-HR annotations into two disjoint crowdsourced batches (fresh workers, same instructions); recompute per-mask prominence and measure Spearman rank correlation between batches. If ρ < 0.65 on >30 % of masks, the evaluation targets are too unstable for the metric audit to be conclusive.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that SSIM/DISTS yield strong localized prominence signals while others fail—treats the crowdsourced fractions (prominence = fraction of viewers noticing the artifact) as reliable ground truth. The protocol is described only at high level; no inter-rater agreement (Fleiss’ κ, Krippendorff’s α), bootstrap stability of majority votes, or cross-pool replication is referenced. In the no-ground-truth Urban100-HR subset this is especially acute, because absence of a clean reference could systematically alter what annotators flag. If the labels are noisy or pool-specific, the reported superiority of classical full-reference metrics is not yet demonstrated.","agreement_with_reader":"agree"},"referee_report":{"model":"grok-4.3","summary":"The paper introduces artifact prominence as the fraction of viewers who notice an artifact in a highlighted region of a super-resolved image. It presents a crowdsourced annotation protocol and the SR-Prominence dataset suite (3,935 masks from DeSRA, Open Images, Urban100, and no-ground-truth Urban100-HR), re-annotates DeSRA to find 48.2% of its binary artifacts unnoticed by a majority, audits detectors and metrics, and reports that classical full-reference metrics (especially SSIM and DISTS) yield strong localized prominence signals while no-reference IQA methods and specialized artifact detectors often fail to generalize. The suite is released with an objective scoring protocol for benchmarking without new crowdsourcing.","tokens_in":1921,"tokens_out":522,"duration_ms":25738,"significance":"If the prominence labels are shown to be stable, the dataset and protocol would usefully shift SR artifact evaluation from binary presence to perceptual impact. The reported strength of SSIM and DISTS as localized signals is a concrete, falsifiable observation that could influence metric selection. The public release of the dataset together with a reproducible scoring protocol is a clear strength supporting community benchmarking.","major_comments":[{"comment":"Protocol description (Section 3 / Dataset Construction): no annotation instructions, inter-annotator agreement (Fleiss’ κ, Krippendorff’s α), bootstrap stability of majority votes, or cross-pool replication are reported. These checks are load-bearing for treating the prominence fractions as reliable ground truth, especially for the Urban100-HR no-GT subset.","section":"Section 3"},{"comment":"Metric audit results (Section 5): the claim that SSIM/DISTS provide strong localized signals while other methods fail to generalize rests directly on the crowdsourced prominence values as ground truth. Without the missing agreement and stability statistics, the reported superiority and generalization failures cannot be evaluated for robustness.","section":"Section 5"}],"minor_comments":[{"comment":"Abstract and Section 4: the phrase 'objective scoring protocol' is used without a concrete description or pseudocode; adding a short formal definition would improve clarity.","section":"Abstract"},{"comment":"Dataset release statement: confirming that raw per-annotator votes (not only aggregated prominence) are included would strengthen transparency and allow independent re-analysis.","section":"Dataset release"}],"recommendation":"major_revision","confidential_remarks":null},"author_rebuttal":{"model":"grok-4.3","summary":"We thank the referee for their thoughtful review and for highlighting the importance of validating the crowdsourced protocol. We address each major comment below and commit to revisions that will strengthen the manuscript's claims regarding label reliability and metric evaluation.","responses":[{"response":"We agree these statistics are essential to substantiate the prominence labels as reliable ground truth. In the revised manuscript, we will include the complete annotation instructions in Section 3. We will compute and report inter-annotator agreement metrics including Fleiss’ κ and Krippendorff’s α. We will also add analyses of bootstrap stability for the majority votes and cross-pool replication results. For the Urban100-HR no-ground-truth subset, we will provide dedicated replication details to confirm consistency across annotation pools.","revision_made":"yes","referee_comment":"[Section 3] Protocol description (Section 3 / Dataset Construction): no annotation instructions, inter-annotator agreement (Fleiss’ κ, Krippendorff’s α), bootstrap stability of majority votes, or cross-pool replication are reported. These checks are load-bearing for treating the prominence fractions as reliable ground truth, especially for the Urban100-HR no-GT subset."},{"response":"We acknowledge that the metric audit results in Section 5 depend on the quality of the prominence ground truth. Incorporating the agreement, stability, and replication statistics as described in our response to the Section 3 comment will enable a more rigorous evaluation of the robustness of the findings on SSIM, DISTS, and the generalization failures of other methods. The revised manuscript will include these validations to support the claims.","revision_made":"yes","referee_comment":"[Section 5] Metric audit results (Section 5): the claim that SSIM/DISTS provide strong localized signals while other methods fail to generalize rests directly on the crowdsourced prominence values as ground truth. Without the missing agreement and stability statistics, the reported superiority and generalization failures cannot be evaluated for robustness."}],"tokens_in":1492,"tokens_out":431,"duration_ms":25728,"standing_objections":[]},"desk_editor":{"model":"grok-4.3","letter":"The one or two things to know: this work defines prominence as the fraction of viewers who notice a highlighted artifact and releases a dataset suite to measure it for super-resolution outputs. It reports that classical full-reference metrics, especially SSIM and DISTS, give stronger localized signals than no-reference IQA or specialized detectors.\n\nThey collected 3935 masks from DeSRA, Open Images, Urban100, and a no-ground-truth Urban100-HR set. Re-annotating DeSRA found 48.2% of its binary labels were not noticed by a majority of viewers. The release includes an objective scoring protocol so new metrics can be tested without fresh crowdsourcing.\n\nThe soft spot is the missing validation for those labels. The abstract supplies no inter-annotator agreement figures, no bootstrap or cross-pool stability tests, and no details on how prominence was aggregated. This gap is largest for the no-ground-truth images, where the lack of a clean reference could shift what annotators flag. If the majority votes are noisy or pool-specific, the metric comparisons rest on shaky ground.\n\nThe work is aimed at SR and perceptual IQA researchers who want evaluation tied to actual viewer impact. A reader building or benchmarking metrics would get a concrete suite to try, provided the labels hold up.\n\nIt shows honest engagement with the limits of binary detection. Send it to peer review so the annotation protocol and statistics can be examined in full.","headline":"The paper defines artifact prominence via crowdsourcing and finds SSIM/DISTS track it better than other methods, but the labels lack any reported stability checks.","tokens_in":2424,"tokens_out":369,"would_cite":false,"duration_ms":26111,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"grok-4.3","headline":"Classical full-reference metrics like SSIM and DISTS predict perceptual prominence of super-resolution artifacts more reliably than no-reference methods or specialized detectors.","keywords":["super-resolution","artifact prominence","crowdsourced evaluation","perceptual image quality","full-reference metrics","SSIM","DISTS","artifact detection"],"falsifier":"A fresh crowdsourced annotation round on the same image regions with a different viewer pool produces prominence values that diverge substantially from the original labels and from the predictions of SSIM or DISTS.","tokens_in":2699,"feed_emoji":"🖼️","tokens_out":730,"duration_ms":24968,"temperature":0.7,"pith_summary":"The paper defines artifact prominence as the fraction of viewers who notice a defect in a highlighted region of a super-resolved image. It releases a crowdsourced dataset suite of 3,935 such annotated masks drawn from multiple sources, including a realistic setting without ground-truth references. Tests across the suite show that standard full-reference metrics supply localized signals that match human judgments of noticeability, while many other approaches do not hold up when the dataset or reference condition changes. The release includes an objective scoring protocol that lets new metrics be checked against the prominence labels without repeating human annotation. This changes evaluation from counting defects to weighing their visible effect on viewers.","feed_headline":"SSIM and DISTS predict where SR artifacts catch the eye","feed_subtitle":"Crowdsourced labels across multiple datasets show full-reference metrics align with viewer judgments better than no-reference alternatives.","key_machinery":"artifact prominence, defined as the fraction of viewers who judge a highlighted region to contain a noticeable artifact, used as the target measure for perceptual impact instead of binary defect presence","core_discovery":"Artifact prominence is defined as the fraction of viewers who judge a highlighted region to contain a noticeable artifact. A crowdsourced protocol produces the SR-Prominence dataset suite with 3,935 masks from DeSRA, Open Images, Urban100, and a no-ground-truth Urban100-HR setting. Re-annotation of DeSRA shows 48.2 percent of its prior binary artifacts are not noticed by a majority of viewers. Classical full-reference metrics, especially SSIM and DISTS, provide strong localized prominence signals while no-reference IQA methods and specialized artifact detectors fail to generalize across datasets and reference settings.","pith_inferences":["Training objectives for future super-resolution networks could use prominence estimates from SSIM or DISTS to suppress only the defects that most viewers notice.","The crowdsourced protocol could be applied to measure perceptual impact in related tasks such as image denoising or compression artifact evaluation.","The consistent failure of no-reference methods points to a need for metrics that better capture localized viewer attention without a clean reference image."],"forward_implications":["48.2 percent of binary artifacts previously labeled in DeSRA are not noticed by a majority of viewers under the new protocol.","SSIM and DISTS can supply localized predictions of where artifacts will stand out to people.","The released objective scoring protocol lets any new metric be benchmarked on the suite without further crowdsourcing.","Super-resolution methods can be compared on the basis of perceptual impact rather than binary defect counts."],"fun_headline_variants":["SSIM and DISTS best predict SR artifact visibility","SR-Prominence dataset benchmarks perceptual artifact metrics","Full-reference metrics outperform in SR artifact evaluation","No-reference IQA lags on SR artifact prominence signals"],"cache_read_input_tokens":2112,"weakest_assumption_plain":"The crowdsourced majority-vote prominence values remain stable when applied to new annotator pools and image sources beyond those used in the study.","fun_headline_variants_meta":{"raw":{"variants":["SSIM and DISTS best predict SR artifact visibility","SR-Prominence dataset benchmarks perceptual artifact metrics","Full-reference metrics outperform in SR artifact evaluation","No-reference IQA lags on SR artifact prominence signals"]},"model":"grok-4.3","cost_usd":0.00505,"raw_usage":{"total_tokens":2516,"prompt_tokens":777,"num_sources_used":0,"completion_tokens":44,"cost_in_usd_ticks":50499500,"prompt_tokens_details":{"text_tokens":777,"audio_tokens":0,"image_tokens":0,"cached_tokens":256},"completion_tokens_details":{"audio_tokens":0,"reasoning_tokens":1695,"accepted_prediction_tokens":0,"rejected_prediction_tokens":0}},"tokens_in":777,"tokens_out":44,"duration_ms":15009,"temperature":1.0,"reasoning_tokens":1695,"cache_read_input_tokens":256,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-06-30T21:28:01.208705+00:00","model_set":{"reader":"grok-4.3"},"falsifier":"A fresh crowdsourced annotation round on the same image regions with a different viewer pool produces prominence values that diverge substantially from the original labels and from the predictions of SSIM or DISTS.","supporting_citations":[],"review_version":1}