{"id":"31da7a6c-f4f1-4df4-9da5-77ff276cc2ae","arxiv_id":"2506.22790","paper_version":2,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":5.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":4,"one_line_summary":"This challenge report shows the top no-reference model reaches SROCC 0.945 on HDR/SDR video, outperforming VMAF, though the benchmark uses a small test set and a narrow baseline set.","lead":"A video-quality challenge tested seven AI models on how well they judge both HDR and standard definition video quality. The winning model, a vision-language AI, beat the industry standard VMAF, but the comparison lacks statistical rigor and covers only a small test set.","discovery_kind":"extension","skeptic_critique":{"model":"deepseek-v4-flash","headline":"The state-of-the-art claim rests on SROCC differences that may not be statistically significant given the 198-video test set clustered into only 11 contents; no confidence intervals or significance tests are reported.","rationale":"I read the paper as a transparent challenge report: it describes the dataset, evaluation protocol, all submitted models, and openly analyzes the SDR/HDR recovery issue in Section V. The release of the dataset and code is a genuine contribution. However, the central claim — that the top-performing model sets a new state-of-the-art benchmark — depends on the reliability of the performance ranking. The reader identified the SDR/HDR recovery protocol as the weakest assumption. I agree that protocol choice is a significant issue, but Section V shows SLCV remains top (0.96) under the joint protocol, so the protocol concern does not destabilize the specific claim about SLCV. The more fundamental problem is that the reported SROCC differences are small relative to the expected sampling variability given only 11 content clusters and no significance testing. The paper never reports confidence intervals, bootstrap estimates, or any statistical test. This is a load-bearing gap because 'top-performing' and 'state-of-the-art' are inherently comparative and probabilistic claims. A concrete content-clustered bootstrap would settle whether the ranking is real or noise. Since the reader already flagged the absence of significance testing in the rationale, my concern reinforces the conditional verdict rather than changing it. I therefore recommend keeping the verdict conditional (UNCHANGED), with a concrete request for clustered confidence intervals or a significance analysis before the headline claim is accepted as stated.","tokens_in":9469,"tokens_out":4917,"duration_ms":56403,"concrete_test":"Perform a content-clustered bootstrap: resample the 11 test contents with replacement (e.g., 10,000 iterations), compute SROCC for SLCV, EQM, P1204.3, VMAF, and the other models on the full set of sequences corresponding to the resampled contents, and report 95% percentile bootstrap confidence intervals and pairwise difference distributions. If the confidence interval for SLCV's SROCC overlaps with those of EQM or P1204.3, the top-performing claim is not statistically supported. Also compute the SROCC separately within each content to assess variability across the 11 contents.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim that SLCV achieves state-of-the-art performance (SROCC 0.945 vs. 0.925 for EQM and P1204.3) is not supported by any significance testing. The test set contains 198 sequences but only 11 distinct source contents, each with 18 distorted versions. SROCC computed over these 198 sequences is not an independent sample; the effective sample size is closer to 11 content clusters. Under a naive independence assumption, the standard error of SROCC around 0.945 with n=198 is roughly sqrt((1-0.945^2)/197) ≈ 0.018, so the 0.02 gap between SLCV and the next-best baselines is about one standard error. When accounting for content clustering, the uncertainty is substantially larger. Thus, the apparent ranking of SLCV over EQM and P1204.3 could easily arise from chance. The paper's own Section V analysis shows that SLCV's SROCC actually improves to 0.96 under the joint SDR/HDR recovery protocol, while EQM, P1204.3, and VMAF drop sharply, so the reader's protocol concern does not directly threaten SLCV's top position. The more load-bearing weakness is that the reported margin of victory is within the noise floor of the evaluation. Additionally, the test set is drawn from the same dataset and encoding pipeline as the training set, so 'generalizable' is only demonstrated within this narrow distribution. Without cross-dataset validation or content-clustered confidence intervals, the headline claim of a new state-of-the-art benchmark is not established.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper reports the design, results, and analysis of the ICME 2025 Grand Challenge on Generalizable HDR and SDR Video Quality Measurement. It describes the HDRSDR-VQA test data, the evaluation protocol based on SROCC/PLCC/KROCC/RMSE after nonlinear mapping, the seven submitted models across FR and NR tracks, and the comparison against P.1204.3, EQM, VMAF, and PSNR-Y. The reported top performer is the NR model SLCV (SROCC 0.945, PLCC 0.943), with four of seven submissions beating VMAF. The paper also includes an analysis showing that performance changes substantially when SDR and HDR scores are recovered jointly rather than separately.","tokens_in":9848,"tokens_out":3611,"duration_ms":34918,"significance":"If the reported results hold, the challenge provides a useful public benchmark and dataset for HDR/SDR VQA, and the SLCV result is a notable demonstration that a fine-tuned multimodal LLM with HDR-to-SDR preprocessing can rank compressed HDR/SDR videos competitively with dedicated metrics. The submission pipeline (Docker containers, held-out test set, released code and data) is a concrete strength that supports reproducibility. However, the headline 'state-of-the-art' claim rests on small SROCC gaps without statistical validation and on a single scoring protocol, so the significance is conditional.","major_comments":[{"comment":"The claim that SLCV is state-of-the-art rests on SROCC differences (0.945 vs. 0.925 for EQM and P1204.3) that are not accompanied by confidence intervals or significance tests. Since the 198 test sequences come from only 11 source contents, the effective number of independent observations is small; content-clustered bootstrap or permutation tests are needed before the ranking can be considered established.","section":"Section III, Table I"},{"comment":"The official evaluation uses the scores in which SDR and HDR are rated as separate content, and the paper reports that under joint recovery SLCV improves to 0.96 while EQM drops from 0.92 to 0.68, P1204.3 from 0.92 to 0.62, and VMAF from 0.90 to 0.55. Because the challenge's stated goal is generalization across HDR and SDR, the protocol choice is load-bearing: different plausible scoring protocols materially change the ranking of methods. Please justify the choice (or report both), and discuss how the 'generalizable' claim depends on it.","section":"Section V"},{"comment":"The test set is drawn from the same HDRSDR-VQA dataset and the same encoding pipeline as the training set, so the reported 'generalizable' performance is only demonstrated within this distribution. No cross-dataset or domain-shift evaluation is provided; without it, the title claim of generalizability is not fully supported.","section":"Section II and Section III"}],"minor_comments":[{"comment":"The heading 'cvvdpMITransformer' appears to be a typo for 'cvvdpMlTransformer' (and similarly 'cvvdpMISaliency' should be 'cvvdpMlSaliency').","section":"Section IV-C"},{"comment":"The footnote states that the challenge is sponsored by Amazon Prime Video, but no sponsor role or conflict-of-interest statement is provided; please clarify the sponsor's involvement.","section":"Section II"},{"comment":"The phrase 'For simplicity' does not justify the choice of scoring protocol; consider replacing it with a principled rationale or a reference to the companion dataset paper.","section":"Section V"},{"comment":"The inference-time footnote says 'unless noted otherwise,' but the CPU/GPU distinction is only in the text; label the table to indicate which rows used CPU vs. GPU.","section":"Table I"},{"comment":"The abstract claims 'state-of-the-art performance,' but the paper only compares with the challenge baselines and submissions; consider qualifying to 'among the methods evaluated here.'","section":"Abstract and Section III"}],"recommendation":"major_revision","confidential_remarks":"The winning teams are co-authors of this challenge summary. This is a common format for grand-challenge papers, but the editor should ensure the review process is aware of the potential conflict and that the held-out evaluation mitigates self-selection bias. No further concerns about novelty disclosure."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"This is a competent challenge report, and the genuinely new thing is the benchmark itself: a curated subset of HDRSDR-VQA, a fixed train/test split, Dockerized submissions, and the first head-to-head comparison of seven models plus four baselines on that split. The authors also released code and training data, and they deserve credit for including Section V, which examines how the subjective recovery protocol (SDR and HDR scored separately vs. jointly) changes results. That kind of self-critical analysis is rare and useful.\n\nThe soft spots are real but not fatal. The headline claim that the top model sets a new state-of-the-art is not supported by the evidence. Only four baselines were used, no confidence intervals or significance tests are reported, and the test set is 198 sequences drawn from just 11 source contents. With clustering by content, the SROCC gap between SLCV (0.945) and EQM or P1204.3 (0.925) is well within plausible sampling noise. So the ranking of the top few methods should be treated as indicative, not definitive. The 'generalizable' in the title is also doing more work than the evidence allows, since the test set comes from the same dataset and encoding pipeline as the training set; cross-dataset validation would be needed to support that word.\n\nOne concern sometimes raised about this paper is the protocol sensitivity shown in Section V. I think that concern is real but does not actually threaten the top position: under the joint recovery protocol, SLCV's SROCC improves to 0.96 while EQM, P1204.3, and VMAF drop sharply. So the winner is robust to that choice, even though the rest of the ranking is not. The bigger issue is purely statistical.\n\nThe author-participant overlap is a conflict of interest, but the evaluation is against held-out test data, so the ranking is not circular by construction. I would not call this a fatal flaw, though independent verification would help.\n\nWho is this for? VQA practitioners who want a reproducible benchmark on compressed HDR/SDR video. It deserves a serious referee, but only with revisions: add content-clustered bootstrap confidence intervals, soften the state-of-the-art claim to something like 'best among the tested methods under this protocol,' and add at least one cross-dataset check. My recommendation is to send it to peer review with those conditions.","headline":"A solid, transparent challenge report whose 'state-of-the-art' claim outruns its statistics: the winning SROCC margin is within the noise floor of an 11-content test set.","tokens_in":10408,"tokens_out":1709,"would_cite":true,"duration_ms":19733,"reading_group":"maybe","serious_thinker":"yes","would_accept_peer_review":true},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"The top-performing model in the ICME 2025 HDR/SDR VQA challenge was a no-reference vision-language model that beat VMAF and every other baseline.","keywords":["video quality assessment","HDR","SDR","no-reference","full-reference","multimodal large language model","grand challenge","JOD scores"],"falsifier":"Re-run the official test evaluation but recover the ground-truth JOD scores with SDR and HDR clips of the same source rated together in the same pairwise sessions (same-source recovery), then recompute SROCC for all models. If SLCV no longer beats P1204.3 or VMAF under that protocol, the state-of-the-art result holds only for the separate-source scoring choice.","tokens_in":9299,"feed_emoji":"🎬","tokens_out":8506,"duration_ms":76005,"temperature":0.7,"pith_summary":"This paper reports the outcomes of the ICME 2025 Grand Challenge on joint HDR and SDR video quality measurement. The organizers built a test bed from 31 open-source 8K HDR sources, each encoded in both HDR10 and SDR at nine bitrate-resolution levels, with subjective scores collected by pairwise comparisons and scaled to JOD units. Against this benchmark, the top no-reference model—a vision-language model fine-tuned with LoRA and fed HDR frames through an HDR-to-SDR preprocessing step—achieved an SROCC of 0.945 and PLCC of 0.943, outperforming the full-reference VMAF baseline (0.905) and several other submitted models. The paper's central claim is that a single dynamic-range-agnostic MLLM pipeline can rank compressed HDR and SDR video quality more accurately than established metrics, provided the evaluation protocol treats SDR and HDR as separate content when recovering subjective scores. That last condition matters: the paper itself shows that if SDR and HDR are scored jointly as the same source, most models' correlations drop sharply.","feed_headline":"Fine-tuned AI model ranks HDR and SDR video quality better than VMAF","feed_subtitle":"Beats full-reference VMAF with a 0.945 SROCC, setting a new mixed HDR/SDR quality benchmark.","key_machinery":"The load-bearing machinery is the evaluation protocol paired with the winning model's preprocessing. The protocol derives ground-truth scores by converting pairwise human comparisons into continuous JOD values, and it recovers SDR and HDR scores as separate content; this separate-source recovery is what makes the reported rankings possible. On the model side, the key object is a multimodal large language model (an InternVL 2.5 vision-language model fine-tuned with LoRA), combined with FFmpeg Mobius tone mapping of HDR frames to SDR, linearization at 1000 nits, and sliding-window spatial crops of two-thirds of the shortest side. The preprocessing harmonizes HDR and SDR inputs into one distribution; the crops multiply the 360-video training set; and averaging six window predictions produces the final score.","core_discovery":"The central discovery is a benchmark result: in a standardized comparison on unseen HDR10 and SDR test content, the no-reference SLCV model reached an SROCC of 0.945 and PLCC of 0.943 across the full test set, with per-format scores of 0.933/0.953 SROCC for SDR and HDR10 respectively. This beats every baseline the organizers ran, including the standardized no-reference model P.1204.3 (0.925), the encoding-feature metric EQM (0.925), and the widely used full-reference VMAF (0.905), and it also edges out the best full-reference submissions. Four of the seven submitted models outperformed VMAF. The authors attribute the win to a recipe: convert HDR to SDR with Mobius tone mapping under a 1000-nit reference display assumption, sample spatial windows at two-thirds of the short side, and fine-tune InternVL 2.5 with LoRA for four epochs, averaging per-window predictions. The paper also documents that the ranking is protocol-dependent: when SDR and HDR versions of the same source are recovered jointly as one content in the subjective analysis, the SROCC of EQM falls from 0.92 to 0.68, P1204.3 from 0.92 to 0.62, VMAF from 0.90 to 0.55, while SLCV's correlation barely changes (0.94 to 0.96).","pith_inferences":["The near-flat SLCV correlation under joint recovery (0.94 to 0.96) suggests the MLLM may have implicitly learned HDR-versus-SDR preferences, not just within-format quality; a targeted test would be to fine-tune the same pipeline on jointly recovered labels and see if it can still rank cross-format pairs.","Because SLCV takes about 42 seconds per video on an L40S while EQM takes about 17 seconds, the performance gap may come with a deployment cost; a practical follow-up would distill the MLLM into a lightweight regressor or reduce the number of spatial windows.","The sharp VMAF drop under joint recovery suggests VMAF is not merely less accurate but miscalibrated across dynamic ranges; adding a dynamic-range-aware calibration layer could narrow the gap without a new architecture.","The challenge used a single encoder (libx265) and a single HDR-to-SDR conversion operator; testing the winning recipe on other codecs and tone-mapping operators would show whether the generalization claim extends beyond the benchmark's distortions."],"forward_implications":["A no-reference model can outperform the full-reference VMAF on compressed HDR/SDR content, suggesting that reference signals may be unnecessary for this rating task.","The same MLLM pipeline, with tone-mapped inputs, can serve both HDR and SDR, which is a practical recipe for streaming services that must monitor mixed-format libraries.","The released benchmark gives future VQA models a fixed protocol and dataset, so new claims of generalization across dynamic ranges can be tested against the same numbers.","Because the protocol choice changes model rankings, challenge reports should report both separate-source and joint-recovery scores; the separate-source numbers alone may overstate cross-format generalization."],"supporting_citations":[{"why":"Supplies the HDRSDR-VQA database subset used for training and test, including the JOD labels.","marker":"[1]"},{"why":"Provides the 31 open-source 8K HDR source videos from which SDR and HDR sequences were derived.","marker":"[2]"},{"why":"The ASAP algorithm that selected the pairwise comparisons used to collect subjective scores.","marker":"[4]"},{"why":"The pwcmp software that converted pairwise comparison data into continuous JOD scores.","marker":"[5]"},{"why":"The winning no-reference model submission whose SLCV results are the paper's headline claim.","marker":"[7]"},{"why":"The standardized no-reference baseline P.1204.3 that the winning model is compared against.","marker":"[11]"},{"why":"The EQM baseline, an encoding-feature no-reference metric that also exceeds VMAF.","marker":"[12]"},{"why":"The VMAF full-reference baseline that four submitted models outperform.","marker":"[13]"},{"why":"InternVL 2.5, the vision-language backbone that the winning model fine-tunes.","marker":"[14]"},{"why":"LoRA, the low-rank adaptation technique used to fine-tune the backbone on 360 videos.","marker":"[15]"}],"fun_headline_variants":["No-reference model beats full-reference VMAF on HDR/SDR video","Fine-tuned AI metric tops VMAF for HDR and SDR quality","SLCV model wins ICME challenge, outperforms VMAF","InternVL VQA scores 0.945 SROCC, surpassing VMAF","Mixed HDR/SDR quality: one no-reference model beats them all"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The rankings hinge on the choice to recover SDR and HDR subjective scores as separate content; if the two formats are treated as the same source in the pairwise analysis, most models' correlations drop sharply and the reported winner may change.","fun_headline_variants_meta":{"raw":{"variants":["No-reference model beats full-reference VMAF on HDR/SDR video","Fine-tuned AI metric tops VMAF for HDR and SDR quality","SLCV model wins ICME challenge, outperforms VMAF","InternVL VQA scores 0.945 SROCC, surpassing VMAF","Mixed HDR/SDR quality: one no-reference model beats them all"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.00032,"raw_usage":{"total_tokens":1848,"prompt_tokens":1032,"completion_tokens":816,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":648,"completion_tokens_details":{"reasoning_tokens":714}},"tokens_in":648,"tokens_out":816,"duration_ms":48952,"temperature":1.0,"reasoning_tokens":714,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-06T21:56:53.551463+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Re-run the official test evaluation but recover the ground-truth JOD scores with SDR and HDR clips of the same source rated together in the same pairwise sessions (same-source recovery), then recompute SROCC for all models. If SLCV no longer beats P1204.3 or VMAF under that protocol, the state-of-the-art result holds only for the separate-source scoring choice.","supporting_citations":[{"cited_title":"Hdrsdr-vqa: A subjective video quality dataset for hdr and sdr comparative evaluation,","cited_arxiv_id":null,"evidence_quote":"Supplies the HDRSDR-VQA database subset used for training and test, including the JOD labels."},{"cited_title":"Avt-vqdb-uhd-2-hdr: An open 8k hdr source dataset for video quality research,","cited_arxiv_id":null,"evidence_quote":"Provides the 31 open-source 8K HDR source videos from which SDR and HDR sequences were derived."},{"cited_title":"Active sampling for pairwise comparisons via approximate message passing and information gain maximization,","cited_arxiv_id":null,"evidence_quote":"The ASAP algorithm that selected the pairwise comparisons used to collect subjective scores."},{"cited_title":"Multimodal large language model for hdr & sdr video quality measurement,","cited_arxiv_id":null,"evidence_quote":"The winning no-reference model submission whose SLCV results are the paper's headline claim."},{"cited_title":"Bitstream-based model standard for 4k/uhd: Itu-t p.1204.3 – model details, evaluation, analysis and open source implementation,","cited_arxiv_id":null,"evidence_quote":"The standardized no-reference baseline P.1204.3 that the winning model is compared against."},{"cited_title":"Encoder-quantization-motion-based video quality metrics,","cited_arxiv_id":null,"evidence_quote":"The EQM baseline, an encoding-feature no-reference metric that also exceeds VMAF."},{"cited_title":"Vmaf: The journey continues,","cited_arxiv_id":null,"evidence_quote":"The VMAF full-reference baseline that four submitted models outperform."},{"cited_title":"Expanding performance boundaries of open-source multimodal models with model, data, and test-time scaling,","cited_arxiv_id":null,"evidence_quote":"InternVL 2.5, the vision-language backbone that the winning model fine-tunes."},{"cited_title":"Lora: Low-rank adaptation of large language models,","cited_arxiv_id":null,"evidence_quote":"LoRA, the low-rank adaptation technique used to fine-tune the backbone on 360 videos."}],"review_version":1}