{"id":"9cc5f410-ce11-40e8-a299-e7bbbe463ee7","arxiv_id":"2506.12269","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":1,"one_line_summary":"Under causal low-delay video conferencing SR, single-image baselines won the general and talking-head tracks, OCR-based text recovery won the screen-content track, and PSNR/SSIM correlated weakly with subjective quality.","lead":"This report describes the ICME 2025 video super-resolution challenge for video conferencing, in which five teams and several baselines were compared under a causal low-delay constraint with H.265-compressed inputs. It releases a new screen-content dataset and finds that human quality ratings and common objective metrics disagree for general and talking-head video, while OCR-based text readability aligns with subjective scores on screen content.","discovery_kind":"new_application","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Model-level correlations over 9-10 models lack confidence intervals and span a narrow quality range, so the claim that objective metrics are insufficient for ranking is not statistically supported.","rationale":"The reader's weakest assumption focuses on the reliability and statistical power of the crowdsourced subjective test and the small model-level sample. I agree: the central claim about low objective-subjective correlation in Tracks 1 and 2 is the most load-bearing part of the paper, and it depends on correlation coefficients computed from only 9-10 models with no uncertainty quantification. The narrow range of CMOS values further attenuates correlations and makes the coefficients sensitive to any single model. This does not necessarily invalidate the paper, but it means the general recommendation to prefer subjective evaluation is not strongly supported by the reported evidence. I also noted an additional concrete issue in Track 3: Equation 1 as typeset, score = CMOS/(-3) + CER/2, does not reproduce the ranking in Table II; a formula such as (CMOS+3)/6 + (1-CER)/2 does match. However, the qualitative conclusion that participants outperformed baselines in screen content remains robust to the formula correction, so the Track 3 claim is less load-bearing. The verdict should remain CONDITIONAL, requiring additional statistical reporting and a corrected equation before acceptance.","tokens_in":9683,"tokens_out":12140,"duration_ms":138561,"concrete_test":"Using the per-clip CMOS and objective metrics that the organizers must have (the paper reports only model-level aggregates), recompute the correlations in Tables IV and V at the clip level (e.g., 20 test clips × 9-10 models) and bootstrap 95% CIs over clips and models. If the clip-level Pearson for PSNR/SSIM/VMAF is substantially higher than the model-level values (e.g., >0.3 higher) or if the bootstrap CI for the model-level VMAF coefficient includes 0, the 'low correlation' conclusion is an artifact of aggregating a small, narrow-range sample and should be qualified.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The paper's central analytical claim—that objective metrics correlate weakly with subjective quality and are therefore inadequate for ranking in Tracks 1 and 2—rests on model-level correlation coefficients (Tables IV and V) computed over only 9 models (Track 1) or 10 models (Track 2), without confidence intervals, significance tests, or leave-one-out sensitivity analysis. The quality range is narrow: Track 1 CMOS spans roughly -2.2 to -2.7 and all models are clustered within about 0.5 MOS units, which restricts the correlation and can attenuate it. For example, the VMAF Pearson coefficient of 0.581 has a 95% bootstrap CI that includes 0, and LPIPS's strong Pearson (-0.893) is driven largely by the single lowest-CMOS model; removing one model could move these coefficients by more than 0.2. The paper uses these fragile coefficients to recommend subjective evaluation for ranking, but the evidence does not establish that the low correlations generalize beyond this small, narrow-range model set.","agreement_with_reader":"agree"},"referee_report":{"model":"deepseek-v4-flash","summary":"The manuscript reports on the ICME 2025 Grand Challenge on Video Super-Resolution for Video Conferencing, covering three tracks (general-purpose, talking head, screen content). It describes the datasets, the causal low-delay setup, the crowdsourced CCR subjective evaluation, and the results, including baseline and participant models. The main findings are that objective metrics (PSNR, SSIM, VMAF, LPIPS) correlate only weakly with subjective CMOS for general-purpose and talking-head content, while for screen content the top-performing teams used OCR- and CER-aware approaches. The paper also open-sources a new screen-content dataset and an extension of the VCD talking-head dataset.","tokens_in":9850,"tokens_out":10644,"duration_ms":106290,"significance":"If the correlation finding is statistically robust, the paper would provide practical evidence that common objective metrics mislead model ranking for video-conferencing SR, and that screen-content SR is best evaluated with text-recovery measures. The release of the datasets is a valuable contribution. The paper also includes a sensible test-retest control for the CER measure (0.0024) and transparently discloses that one baseline violates the causal constraint. However, the correlation claim is currently supported only by small-sample point estimates without confidence intervals, and the subjective scoring procedure lacks rater-level statistics, so the significance of the headline finding is not yet established.","major_comments":[{"comment":"The central claim that objective metrics correlate weakly with subjective quality in Tracks 1 and 2 rests on model-level correlation coefficients computed from only 9 models (Track 1) and 10 models (Tracks 2 and 3). No confidence intervals, significance tests, or leave-one-out sensitivity analyses are provided. Given the narrow CMOS range (e.g., -2.216 to -2.712 in Track 1), the point estimates are fragile; for instance, the Track 1 VMAF Pearson coefficient of 0.581 could plausibly include zero, and the LPIPS Pearson coefficient of -0.893 may be driven by the single lowest-CMOS model. The authors should supply bootstrap (or jackknife) confidence intervals for all coefficients in Tables IV and V, or at minimum temper the conclusion to acknowledge that the low correlations are not established beyond this small, narrow-range model set.","section":"III-A, Tables IV and V"},{"comment":"The reliability of the subjective scores is not documented. The paper reports per-model CMOS confidence intervals in Table II but no rater counts, no inter-rater agreement (e.g., ICC or Krippendorff's alpha), and no per-clip confidence values. This matters because the tie-ranking procedure in Section III depends on 'no significant difference between the distributions of CMOS values,' and because the correlation analysis in Section III-A relies on the precision of the CMOS values. Add the number of raters per clip, the rating design, and an inter-rater reliability statistic.","section":"II-B"},{"comment":"The rule for assigning tied ranks in Table II is not operationalized. The text states that two consecutive models are tied when 'there is no significant difference between the distributions of CMOS values,' but no statistical test is named, and no threshold is given. If the method is the CI-overlap approach described in Ref. [28], state that explicitly and describe how the 95% CIs are used. Without this, the ranks, on which several qualitative conclusions depend, are not reproducible.","section":"III"},{"comment":"The Track 3 challenge score in Eq. (1) is ambiguous as typeset. The text calls it a 'normalized average of CMOS and CER,' but the equation can be read either as (CMOS/-3) + (CER/2) or as CMOS/(-3 + CER/2). The ranking in Table II is consistent with the latter interpretation, but the equation should be typeset as a single unambiguous fraction (e.g., score = CMOS / (-3 + CER/2)) so that the scoring rule is reproducible from the text alone.","section":"II-B, Eq. (1)"}],"minor_comments":[{"comment":"The Track 3 training-set counts (1520 GT, 9072 LR+H.265) do not match the expected six QP levels per clip (1520x6=9120). Please clarify whether this is due to encoding failures or to a different number of QPs for some clips.","section":"II-A, Table I"},{"comment":"The specific H.265 QP values used are not listed. Providing the QP numbers would improve reproducibility.","section":"II-A"},{"comment":"The CER measurement pipeline is under-specified; the OCR engine and the frame/text-region selection criteria should be stated, as CER is a key outcome for Track 3.","section":"II-B"},{"comment":"The phrase 'Spearman correlation coefficient of ≤0.8' is confusing; in Track 1 the value is -0.800, so it should be phrased as 'an absolute value no greater than 0.8.'","section":"III-A, Table IV"},{"comment":"The statement that participating teams 'significantly outperformed' baselines in Track 3 is not backed by a significance test; the non-overlapping CIs in Table II are suggestive, but the paper should either perform a formal test or refer to the CI-based tie-breaking rule.","section":"III"},{"comment":"The FLOPs value for BVIVSR appears twice as 455.16; please verify whether the FLOPs are identical for Track 1 and Track 2, and fix the formatting of the table.","section":"Table VI"}],"recommendation":"major_revision","confidential_remarks":"The paper is a challenge-report and dataset-release manuscript; its fit for a journal depends on the editorial scope. The authors should also disclose the potential conflict of interest that the subjective-test implementation [23] and the VCD dataset [22] are their own works; this is not circular scientifically, but it should be acknowledged for transparency."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"The short version: this is a solid, honest challenge report that earns its keep through the new open screen-content dataset and the careful evaluation setup. The qualitative findings—baselines beat participants on natural video, text-aware screen-content models win their track, and PSNR/SSIM rank models poorly for conference content—are plausible and consistent with prior work. The paper deserves a serious referee, but the correlation analysis needs work before publication.\n\nWhat is genuinely new: a 95-clip lossless screen-content dataset with 9,072 training LR clips, a causal low-delay evaluation protocol (no future frames), and the use of OCR-derived CER alongside standard subjective CMOS. The test-retest CER baseline of 0.0024 is a good control, and the authors transparently disclose that two video baselines violate the causal constraint. Tables are internally consistent, and the team descriptions are informative.\n\nWhere it is soft: the paper's headline claim—that objective metrics correlate weakly with subjective quality and are therefore inadequate for ranking—rests on model-level correlations over only 9–10 models, with no confidence intervals, no leave-one-out sensitivity, and a narrow CMOS range (roughly 0.5 MOS units across all models). The stress-test note is right: a coefficient like VMAF's 0.58 Pearson likely includes zero once uncertainty is accounted for, and removing the lowest-CMOS model could shift LPIPS's strong correlation substantially. The crowdsourced CCR protocol is also reported only at the protocol level; no rater counts, agreement statistics, or per-clip confidence intervals for the CMOS values are given, which matters because those intervals are used to construct tied ranks. Equation 1 is typeset ambiguously; the intended score is presumably (CMOS/−3 + CER/2), but it should be written explicitly. None of these are fatal to the dataset or the qualitative ranking, but they do undercut the strength of the correlation conclusion as stated.\n\nWho this is for: researchers working on video super-resolution, especially for conferencing and screen content, and anyone building benchmarks for perceptual quality. It is a useful reference point, not a theoretical breakthrough.\n\nRecommendation: send it to peer review. Ask the authors to add confidence intervals or bootstrap/leave-one-out analysis for the correlation coefficients, report subjective-test reliability details, fix the equation, and soften the generalizing claim about objective metrics if the robustness analysis stays inconclusive. That is a reasonable revision path, not a reject.","headline":"A useful challenge report with a genuine new dataset, but the central metric-correlation claim is statistically thinner than it looks.","tokens_in":10399,"tokens_out":1417,"would_cite":true,"duration_ms":20695,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The paper argues that for general-purpose and talking-head video super-resolution, PSNR and SSIM correlate weakly with subjective quality, so rankings should rely on subjective evaluation, while screen-content super-resolution is won by…","keywords":["video super-resolution","video conferencing","subjective quality assessment","Comparison Category Rating","CMOS","character error rate","screen content","ITU-T P.910"],"falsifier":"Re-run the subjective evaluation with a larger panel of trained raters on the same test clips and compute the Spearman correlation between CMOS and PSNR or SSIM per track; if in Track 1 or Track 2 that correlation rises above about 0.8, the paper's claim that objective metrics rank poorly would be contradicted. Alternatively, if a held-out set of 50 or more models shows PSNR ranking matching subjective ranking, the result would not generalize.","tokens_in":9490,"feed_emoji":"🎥","tokens_out":6512,"duration_ms":185370,"temperature":0.7,"pith_summary":"This paper reports the organization and results of a grand challenge on video super-resolution for video conferencing, where low-resolution H.265-encoded video must be upscaled under a low-delay, causal constraint. Its central finding is that in the general-purpose and talking-head tracks, standard objective metrics such as PSNR and SSIM correlate weakly with subjective quality scores, so ranking models by these metrics would mislead; subjective evaluation is necessary. In the screen-content track, by contrast, PSNR and SSIM correlate strongly with subjective opinion, and the winning models combined super-resolution with OCR-based text recovery, evaluated using character error rate in addition to subjective scores. The paper also contributes new open datasets, including a screen-content dataset and an extension of the talking-head VCD dataset. A sympathetic reader would care because the result questions common practice in benchmarking video super-resolution models and points to content-dependent evaluation.","feed_headline":"PSNR and SSIM mislead video SR rankings, challenge shows","feed_subtitle":"Objective scores barely track human opinion; screen content hinges on OCR-based CER.","key_machinery":"The central mechanism is the crowdsourced implementation of ITU-T Rec. P.910 Comparison Category Rating (CCR), which produces CMOS scores by having raters compare each processed clip directly against the ground-truth source. For screen content, this is augmented by a composite challenge score that averages normalized CMOS with character error rate (CER), computed from OCR over approximately 210,000 annotated characters. The paper then uses model-level Pearson, Spearman, Kendall's Tau-b, and a confidence-interval-adjusted Tau-b95 to compare these subjective scores against PSNR, SSIM, VMAF, LPIPS, and CER.","core_discovery":"The authors claim that objective metrics and subjective quality scores are only weakly correlated for general-purpose and talking-head video super-resolution, making subjective evaluation necessary for ranking models in these contexts. For Track 1, the Pearson correlation between CMOS and PSNR is -0.212 and with SSIM is -0.073, while LPIPS reaches -0.893 Pearson but only -0.800 Spearman; Track 2 shows a similar pattern, with PSNR Pearson at 0.232, SSIM at 0.316, and LPIPS at -0.732 Pearson and -0.720 Spearman. In contrast, for screen content, PSNR and SSIM show strong correlations with CMOS (Pearson 0.885 and 0.894, respectively), and CER correlates strongly as well. The paper also reports that in the general-purpose and talking-head tracks, pretrained image-based baseline models ranked first in subjective tests, while in the screen-content track, participating teams that used OCR and CER-based methods significantly outperformed all baselines.","pith_inferences":["Since every model's CMOS is negative in all three tracks, none of the submissions improved on the ground-truth source: the practical task was damage minimization, and the challenge measured which model degraded least, a framing the paper does not itself emphasize.","The low correlation in Tracks 1 and 2 may be partly an artifact of the narrow model sample, about ten pipelines per track, many of them off-the-shelf baselines, so the observed coefficients are estimates over a small, nonrandom set.","A natural testable extension is to add a CER-style loss or OCR post-processing to general-purpose and talking-head models, since the screen-content result suggests that training to preserve semantically identifiable structures may also improve perceptual quality."],"forward_implications":["Benchmarks that rank video super-resolution models by PSNR or SSIM on general or talking-head content will not reflect human preference; organizers should budget for subjective testing.","LPIPS is the strongest objective proxy in these tracks, but its Spearman correlation of about 0.8 still leaves enough disagreement that final rankings should be subjective.","For screen content, PSNR and SSIM are usable development metrics, and CER should be reported alongside them whenever text legibility matters.","Screen-content super-resolution systems should include OCR or text-recovery components; the top two teams used them and clearly separated from generic baselines.","The newly released screen-content and extended talking-head datasets allow future training and evaluation outside the challenge."],"supporting_citations":[{"why":"Supplies the crowdsourced CCR method and the CMOS measure that all subjective rankings in the paper rest on.","marker":"[23]"},{"why":"Provides the VCD dataset that Track 2 extends, with the extension being one of the paper's open-sourced contributions.","marker":"[22]"},{"why":"Provides the REDS training and validation clips used for Track 1.","marker":"[20]"},{"why":"Provides the 20 OpenVid-1M test clips used for Track 1.","marker":"[21]"},{"why":"SwinIR, the single-image baseline that achieved first place in Track 2 and serves as a comparison point in all tracks.","marker":"[24]"},{"why":"BasicVSR++, a video baseline adapted to the causal one-frame-at-a-time constraint and used across all tracks.","marker":"[27]"},{"why":"RealSR, the image baseline that achieved first place in Track 1 and is compared across tracks.","marker":"[25]"},{"why":"Defines the confidence-interval-based tie handling used to compute the Tau-b95 correlations between subjective and objective metrics.","marker":"[28]"}],"fun_headline_variants":["PSNR and SSIM mislead video SR rankings, challenge shows","Subjective tests beat PSNR for ranking video SR models","Screen content SR: OCR/CER methods beat image baselines","Image SR baselines win general video; OCR wins screen content"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The conclusions assume the crowdsourced CCR subjective ratings are accurate enough on these 300-frame clips that model-level correlations with about ten models per track are meaningful, even though the paper reports no rater counts, inter-rater agreement, or per-clip confidence.","fun_headline_variants_meta":{"raw":{"variants":["PSNR and SSIM mislead video SR rankings, challenge shows","Subjective tests beat PSNR for ranking video SR models","Screen content SR: OCR/CER methods beat image baselines","Image SR baselines win general video; OCR wins screen content"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.000306,"raw_usage":{"total_tokens":1748,"prompt_tokens":937,"completion_tokens":811,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":553,"completion_tokens_details":{"reasoning_tokens":739}},"tokens_in":553,"tokens_out":811,"duration_ms":9800,"temperature":1.0,"reasoning_tokens":739,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-07T00:54:51.447800+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the subjective evaluation with a larger panel of trained raters on the same test clips and compute the Spearman correlation between CMOS and PSNR or SSIM per track; if in Track 1 or Track 2 that correlation rises above about 0.8, the paper's claim that objective metrics rank poorly would be contradicted. Alternatively, if a held-out set of 50 or more models shows PSNR ranking matching subjective ranking, the result would not generalize.","supporting_citations":[{"cited_title":"A crowdsourcing approach to video quality assessment,","cited_arxiv_id":null,"evidence_quote":"Supplies the crowdsourced CCR method and the CMOS measure that all subjective rankings in the paper rest on."},{"cited_title":"Vcd: A video conferencing dataset for video compression,","cited_arxiv_id":null,"evidence_quote":"Provides the VCD dataset that Track 2 extends, with the extension being one of the paper's open-sourced contributions."},{"cited_title":"Ntire 2019 challenge on video deblurring and super-resolution: Dataset and study,","cited_arxiv_id":null,"evidence_quote":"Provides the REDS training and validation clips used for Track 1."},{"cited_title":"Swinir: Image restoration using swin transformer,","cited_arxiv_id":null,"evidence_quote":"SwinIR, the single-image baseline that achieved first place in Track 2 and serves as a comparison point in all tracks."},{"cited_title":"Basicvsr++: Improving video super-resolution with enhanced propagation and alignment,","cited_arxiv_id":null,"evidence_quote":"BasicVSR++, a video baseline adapted to the causal one-frame-at-a-time constraint and used across all tracks."},{"cited_title":"Real-world super-resolution via kernel estimation and noise injection,","cited_arxiv_id":null,"evidence_quote":"RealSR, the image baseline that achieved first place in Track 1 and is compared across tracks."},{"cited_title":"Transformation of mean opinion scores to avoid misleading of ranked based statistical techniques,","cited_arxiv_id":null,"evidence_quote":"Defines the confidence-interval-based tie handling used to compute the Tau-b95 correlations between subjective and objective metrics."}],"review_version":1}