{"id":"02b962ac-73c8-49ed-9b86-678d7890940f","arxiv_id":"2608.08508","paper_version":1,"verdict":"CONDITIONAL","confidence":"MODERATE","novelty_score":3.0,"correctness_risk":"medium","formal_verification":"none","parameter_count":0,"one_line_summary":"A doctoral summary reports consistent perceptual-quality gains from test-time adaptation for video super-resolution, screen-content super-resolution, and no-reference video quality assessment, using the author's previously published frameworks.","lead":"This paper summarizes three test-time adaptation frameworks the author previously published, which aim to improve video super-resolution and no-reference quality assessment under unknown real-world distortions. A reader may look here for a compact status report on applying test-time adaptation to video enhancement without ground-truth references.","discovery_kind":"review","skeptic_critique":{"model":"deepseek-v4-flash","headline":"Table 2 never names its evaluation metric; if it is the same RW-VSR-QA score used as the loss, the claimed consistent improvements are circular and unverified.","rationale":"The reader's verdict is CONDITIONAL and focuses on the adapted quality model's alignment with human perception. My concern is more specific and prior: the evidence table does not identify its metric, and the loss formulation references HR scores that are unavailable under the stated test-time setting. If Table 2 uses the same RW-VSR-QA score as the loss, the improvements are expected by optimization and prove nothing about perceptual quality. This is an internal-consistency and evidence-completeness issue rather than a disagreement with the consensus. The cited prior publications may supply the missing protocol and independent evaluation, so the conditional verdict remains appropriate: the claim is accepted only if the underlying papers provide a non-circular evaluation. I do not see grounds to reject outright, because the referenced works are real venues and the summarized results may be fully supported there; I also do not see grounds to accept on this manuscript alone, because Table 2's metric is unnamed and the loss's HR-score term is unexplained. No formal verification or code is present to substitute for these missing details.","tokens_in":8025,"tokens_out":4953,"duration_ms":57486,"concrete_test":"Obtain the exact evaluation protocol for Table 2 from reference [20] or the author; then recompute all six backbones on RVSR, MotionBlur, VLQ, and K|Lens using at least two independent no-reference metrics not involved in the loss (e.g., NIQE, MUSIQ, DOVER) on released SR outputs. If the reported gains are absent or reversed under independent metrics, the central claim is a circular artifact; if gains persist, the concern is resolved.","verdict_should_be":"UNCHANGED","load_bearing_attack":"The central claim—that RW-VSR-QA as an adaptive quality loss consistently improves six VSR backbones across four datasets—rests entirely on Table 2, yet the table does not state what metric is being reported. This matters because the same RW-VSR-QA model generates the adaptive quality loss in Section 3.1, Method Part-II. If Table 2 reports RW-VSR-QA's own predicted quality scores, then updating the SR network to minimize the loss against those scores should improve the table by construction, even if human-perceived quality is unchanged or worse. Additionally, the loss is defined as the difference between SR and HR quality scores, but the abstract and Section 2 state that no high-resolution ground truth is available at inference; no procedure for obtaining the HR score is given in this manuscript. The paper cites prior publications [19]–[21] that may contain full protocols and independent metrics, but this manuscript alone does not establish the claim. The temporal-modeling limitation acknowledged in Section 6 is secondary to this missing metric specification and potential circularity.","agreement_with_reader":"partial"},"referee_report":{"model":"deepseek-v4-flash","summary":"The paper, framed as a doctoral research summary, proposes three test-time adaptation (TTA) frameworks for video/image super-resolution and quality assessment: (i) a TTA method for no-reference video quality assessment (RW-VSR-QA) whose adapted quality scores are used as an adaptive loss to guide real-world video super-resolution; (ii) ScrVSR, a transformer-based screen-content video super-resolution model trained with text-aware and perceptual losses; and (iii) SCISR-TTA, a region-aware TTA strategy that adapts screen-content super-resolution models separately on text and non-text regions. The claimed contributions are consistent improvements in perceptual quality and readability across multiple backbones and datasets, with experimental tables supporting each component, and the paper explicitly lists its temporal-modeling limitation as future work.","tokens_in":8187,"tokens_out":3534,"duration_ms":39098,"significance":"If the central claims are correct, the paper would offer a practically valuable recipe for guiding super-resolution at test time without paired high-resolution data, using a TTA-adapted quality model as a loss. The idea is timely and the paper is clearly written in parts, with results reported across several backbones and datasets. The paper also honestly discloses prior publications [19]–[21] for each component and includes an explicit limitations section. However, the evidence presented here is not self-contained: key protocols and metrics are deferred to prior papers, and the main quantitative support (Table 2) does not name its evaluation metric, which creates a risk of circularity because the same quality model supplies both the adaptive loss and, potentially, the reported score. The absence of error bars or significance tests further weakens the strength of the 'consistent improvements' claim.","major_comments":[{"comment":"Table 2 never states the evaluation metric behind the reported numbers or the 'Avg.' column. Since the adaptive quality loss is defined as the difference between SR and HR quality scores produced by the RW-VSR-QA model, if Table 2 reports exactly those scores, then updating the SR network to minimize the loss would improve the table by construction even if human-perceived quality did not improve. The authors should name the metric, state whether it is the RW-VSR-QA score itself or an independent metric, and report at least one metric that is not used as a training objective.","section":"Section 3.1, Method Part-II; Table 2"},{"comment":"The adaptive quality loss is the difference between the quality scores of SR and HR videos, but Section 2 states that no high-resolution reference is available at inference time. The manuscript does not specify how the HR quality score is obtained in the test-time setting; it refers to prior work [20] for that procedure. For the present manuscript to establish the claim on its own, it needs to describe or justify the availability of the HR reference used for the loss.","section":"Section 3.1, Method Part-II; Section 2"},{"comment":"The claim of 'consistent improvements across all backbones and datasets' is quantitatively weak in several entries; for example, HAT_Q improves over HAT by only 0.0001 in Avg., and COVER's SRCC gain in Table 1 is only 0.45%. None of the tables reports error bars, confidence intervals, or significance tests, so it is unclear whether any of the smaller gains are statistically distinguishable from noise. I recommend reporting variance over runs or at least identifying which gains exceed a meaningful threshold.","section":"Table 2; Section 3.1 Results"},{"comment":"The arrow for DISTS is marked as ↑ (higher is better), but the standard definition of DISTS [11] uses lower values to indicate higher similarity. If the table values follow the standard definition, the direction is wrong; if the values have been transformed (for example, negated or inverted), the transformation is not explained. In addition, EIQM decreases for ITSRN_ada (0.1655 to 0.1636) and LIIF_ada (0.2089 to 0.2035) relative to their baselines, so the sentence claiming that SCISR-TTA 'consistently improves both text fidelity and perceptual quality' is not fully supported by the table.","section":"Table 4"},{"comment":"The paper acknowledges that the proposed methods operate on individual frames and do not explicitly model temporal dependencies. Since the headline results are for video super-resolution, the absence of any temporal-consistency metric (for example, flicker or temporal stability) means the claim of improved perceptual quality for videos is incomplete. A frame-wise perceptual improvement does not automatically translate to a better video viewing experience, yet the paper's central 'consistent improvements across all backbones and datasets' claim is made without temporal evaluation.","section":"Section 6; Section 3.1 Results"}],"minor_comments":[{"comment":"Reference [1] contains a truncated author name, 'Del B, A.', and several references contain typographical spacing artifacts such as 'W ang' and 'Y an'; these should be corrected.","section":"References"},{"comment":"The natural-scene loss expression 'L_ns = ||I_ns - Ibar_ns||_2^2' is not numbered or clearly typeset; consider formatting it as a numbered equation.","section":"Section 3.3"},{"comment":"The qualitative comparison in Figure 1 relies on red, blue, and yellow highlighted regions, but the caption does not explain the color coding for all panels; please label the crops directly for readability.","section":"Figure 1"},{"comment":"The sentence 'This work has been published in IEEE Transactions on Artificial Intelligence [20]' appears abruptly at the end of the Results paragraph; consider moving publication notes to a separate 'Publication Note' or footnote.","section":"Section 3.1 Results"},{"comment":"The abstract claims 'consistent improvements in perceptual quality and readability', but Table 4 shows EIQM decreasing for two of the four baselines; please align the abstract's wording with the actual table results or clarify why those decreases are not considered counterexamples.","section":"Abstract; Section 3.3"}],"recommendation":"major_revision","confidential_remarks":"This manuscript reads as an extended summary of three previously published papers ([19], [20], [21]) rather than as a self-contained contribution. The main load-bearing issue is Table 2: the evaluation metric is unnamed and the adaptive loss is produced by the same quality model that may be producing the reported scores, which would make the claimed consistent improvements circular. The authors should be asked to disclose the metric, add an independent evaluation, and either provide the HR-score procedure or temper the claim. Given the reliance on prior publications and the missing protocol details, I would not accept the paper in its current form, but the underlying research direction is valuable and the issues appear fixable within the scope of a revision."},"author_rebuttal":null,"desk_editor":{"model":"deepseek-v4-flash","letter":"Plainly: this is a doctoral research summary, not a new research paper. The three frameworks it describes are already published by the same author in [19], [20], and [21], and the paper says so explicitly. So the novelty is close to zero. What you get is a compact overview with a few tables and one qualitative figure.\n\nTo give credit, the paper is transparent about where each contribution comes from, and the framing of TTA for VSR/VQA is coherent. The screen-content SR area is underserved, and if the prior papers hold up, a summary of this sort could orient new researchers. But as a standalone manuscript, it does not support its own central claim.\n\nThe main issue is Table 2. The caption never says what metric is being reported. The numbers look like a learned score, not PSNR/SSIM. The adaptive quality loss in Method Part-II is defined as the difference between the quality scores of SR and HR videos, but the abstract says no HR ground truth is used, and the paper never explains where the HR score comes from. If Table 2 is reporting the RW-VSR-QA score itself, then the improvement is partly circular—you are optimizing the very score you then present as evidence. The stress-test note is on target.\n\nSecondary issues: no error bars or significance tests, and some gains are trivial (HAT +0.0001). The temporal-modeling limitation is acknowledged, which is honest, but it doubles down on the sense that this is a progress report, not a contribution.\n\nNet: I would not send this to a regular peer-review track. There is no new method, no new result, and the one table that could make a point is not interpretable. If the venue has a doctoral-consortium or extended-abstract track, it might fit with a pointer to the underlying papers. For a research paper, desk-reject is the right call.","headline":"A doctoral research summary that compiles three already-published frameworks; the only new content is a table with an unnamed metric, which makes the central claim unverifiable.","tokens_in":8736,"tokens_out":6137,"would_cite":false,"duration_ms":62959,"reading_group":"no","serious_thinker":"unclear","would_accept_peer_review":false},"rs_alignment":null,"lean_confirmation":null,"pith_extraction":{"msc":[],"pacs":[],"model":"deepseek-v4-flash","headline":"This paper claims that adapting a no-reference quality model to each test batch produces a perceptual loss that consistently improves every super-resolution backbone and dataset it tests, and that region-aware test-time adaptation…","keywords":["Test-Time Adaptation","Video Quality Assessment","Adaptive Quality Loss","Video Super-Resolution","Screen Content Images","Perceptual Quality","Character Error Rate","Contrastive Loss"],"falsifier":"Run the adaptive quality loss on heavily distorted real-world videos that also have human opinion scores, and have viewers choose between base and adapted outputs; if the adapted output is preferred no more often than chance in the clips where the quality model and humans disagree most, the claim that score matching improves perceived quality is refuted.","tokens_in":7793,"feed_emoji":"🎥","tokens_out":9868,"duration_ms":96420,"temperature":0.7,"pith_summary":"Real-world video super-resolution often meets distortions—compression, noise, motion blur—that no training set fully covers, and high-resolution ground truth is usually absent at deployment. The paper's central claim is that a no-reference video-quality model can be adapted on the fly to the statistics of each test video, and the gap between the quality scores it assigns to the super-resolved and reference versions can then be used as an adaptive loss. According to the reported numbers, this adaptive loss improves all six super-resolution backbones on all four real-world degraded datasets, with the largest gains where distortion is most severe. For screen content, the paper further claims that text-aware losses in a transformer-based SR model and a region-aware test-time adaptation step reduce character error rates and improve perceptual metrics without ground-truth supervision at adaptation time. If the claims hold, super-resolution becomes self-adjusting at inference time, a practical path toward deployment on heterogeneous devices and network conditions.","feed_headline":"Adapted quality loss lifts every super-resolution backbone","feed_subtitle":"A no-reference quality model adapted on test batches supplies the perceptual loss, no ground-truth needed.","key_machinery":"The load-bearing mechanism is test-time adaptation of a quality network. On each batch of test video, only the normalization-layer parameters are updated under two auxiliary objectives, Quality-Based Group Contrastive Loss (clips of similar quality are pulled together in the embedding space) and Quality-Aware Rank Loss (predicted quality ordering is kept consistent), so the pretrained model shifts to the current distortion distribution without discarding learned semantics. The adaptive quality loss is then the difference between the quality scores assigned to the super-resolved and reference videos, used as a perceptual regularizer alongside pixel and adversarial losses. For screen content, the machinery is text-aware supervision: a character error rate loss, a CLIP-based quality loss, and a VGG perceptual loss in the supervised transformer model, plus a region-aware two-stage test-time adaptation in which text regions are refined with OCR and small-language-model feedback and non-text regions are refined against a cleaned reference produced by a lightweight restoration network.","core_discovery":"The paper's central discovery, stated on its own terms, is that test-time adaptation makes a no-reference quality model a usable perceptual supervisor for super-resolution. Updating only the normalization layers of a pretrained video quality assessment network, with a quality-based group contrastive loss and a quality-aware rank loss, lets the model track the current distortion mix; the difference between the quality scores of the super-resolved and reference videos then defines an adaptive loss. Across BasicVSR, RBVSR, HAT, SRWD, RESR, and IART, adding this loss improves the reported average quality score on every one of the RVSR, VLQ, K|Lens, and MotionBlur datasets, with the clearest gains where distortion is severe. A second line of results claims that a transformer-based screen-content SR model trained with character error rate, CLIP-based perceptual, and VGG losses beats its baselines in PSNR, SSIM, and CER across quantization levels, and that a region-aware dual-branch test-time adaptation reduces character error and improves perceptual metrics for five screen-content backbones.","pith_inferences":["A natural extension the paper leaves implicit is to use the same score-matching loss for other restoration tasks, such as deblurring or compression-artifact removal, where a test-adapted quality model could replace hand-designed regularizers.","The pattern of gains in Table 1, with the largest gain for the lightweight single-branch VQA model, suggests that the amount of improvement from test-time adaptation may depend on how far the baseline already generalizes; that could be tested by correlating gain size with baseline cross-distortion performance.","The text branch of SCISR-TTA leans on OCR feedback, so a concrete failure test is to break the OCR step and check whether the character error rate gains disappear; the paper does not report this ablation."],"forward_implications":["An existing video super-resolution model can be upgraded at inference time by adding the adaptive quality loss, with no retraining of the SR weights.","No-reference video quality assessment itself becomes more reliable under unseen distortion mixtures after test-time adaptation, as shown by rank-correlation gains on five QA baselines.","Screen-content super-resolution can target readable text and sharp structure rather than only pixel fidelity, with lower character error rates across compression levels.","Region-aware test-time adaptation transfers to backbones not designed for text, reducing character error and improving perceptual scores for all five screen-content baselines tested.","Because the methods operate per frame, temporal consistency and flicker reduction remain open, as the paper itself notes."],"supporting_citations":[{"why":"Supplies the RW-VSR-QA algorithm that this paper uses as the adaptive quality loss for real-world super-resolution.","marker":"[20]"},{"why":"One of the five no-reference VQA baselines whose test-time-adapted version is evaluated in Table 1.","marker":"[23]"},{"why":"A second VQA baseline used in Table 1 to show the adapted quality model improves ranking.","marker":"[24]"},{"why":"A third VQA baseline evaluated in Table 1, providing a comparison point for test-time adaptation gains.","marker":"[12]"},{"why":"The VQA baseline that shows the largest rank-correlation gain after adaptation in Table 1.","marker":"[17]"},{"why":"BasicVSR is one of the six VSR backbones in Table 2 used to demonstrate the adaptive quality loss.","marker":"[2]"},{"why":"The transformer screen-content SR baseline that ScrVSR extends and SCISR-TTA adapts.","marker":"[27]"},{"why":"Supplies the fixed restoration network whose cleaned output is the reference for the non-text region loss in SCISR-TTA.","marker":"[6]"},{"why":"A competing screen-content enhancement method used as comparison in the ScrVSR and SCISR-TTA tables.","marker":"[15]"}],"fun_headline_variants":["Test-time adapted VQA loss lifts all super-resolution backbones","No-reference VQA adapts on the fly to guide super-resolution","Adapted quality model improves every super-resolution backbone","Region-aware TTA sharpens screen content without ground truth","Adapted VQA loss upgrades all SR backbones with no retraining"],"cache_read_input_tokens":3200,"weakest_assumption_plain":"The load-bearing assumption is that the adapted quality model's scores align with human perception on the specific distortions in the test video, so making super-resolved output match those scores actually makes it look better rather than merely matching the model's own biases.","fun_headline_variants_meta":{"raw":{"variants":["Test-time adapted VQA loss lifts all super-resolution backbones","No-reference VQA adapts on the fly to guide super-resolution","Adapted quality model improves every super-resolution backbone","Region-aware TTA sharpens screen content without ground truth","Adapted VQA loss upgrades all SR backbones with no retraining"]},"model":"deepseek-v4-flash","effort":"low","cost_usd":0.001191,"raw_usage":{"total_tokens":4904,"prompt_tokens":924,"completion_tokens":3980,"prompt_tokens_details":{"cached_tokens":384},"prompt_cache_hit_tokens":384,"prompt_cache_miss_tokens":540,"completion_tokens_details":{"reasoning_tokens":3895}},"tokens_in":540,"tokens_out":3980,"duration_ms":33182,"temperature":1.0,"reasoning_tokens":3895,"cache_read_input_tokens":384,"cache_creation_input_tokens":0},"cache_creation_input_tokens":0},"created_at":"2026-08-14T04:32:07.636946+00:00","model_set":{"reader":"deepseek-v4-flash"},"falsifier":"Run the adaptive quality loss on heavily distorted real-world videos that also have human opinion scores, and have viewers choose between base and adapted outputs; if the adapted output is preferred no more often than chance in the clips where the quality model and humans disagree most, the claim that score matching improves perceived quality is refuted.","supporting_citations":[{"cited_title":"K., Mishra, A., Jakhetiya, V., Subudhi, B","cited_arxiv_id":null,"evidence_quote":"Supplies the RW-VSR-QA algorithm that this paper uses as the adaptive quality loss for real-world super-resolution."},{"cited_title":"Fast-vqa: Efficient end-to-end video quality assessment with fragment sampling","cited_arxiv_id":null,"evidence_quote":"One of the five no-reference VQA baselines whose test-time-adapted version is evaluated in Table 1."},{"cited_title":"IEEE Transactions on Pattern Analysis and Machine Intelligence(2023)","cited_arxiv_id":null,"evidence_quote":"A second VQA baseline used in Table 1 to show the adapted quality model improves ranking."},{"cited_title":"InProceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition(2024), pp","cited_arxiv_id":null,"evidence_quote":"A third VQA baseline evaluated in Table 1, providing a comparison point for test-time adaptation gains."},{"cited_title":"InProceedings of the AAAI Conference on Artificial Intelligence(2024), vol","cited_arxiv_id":null,"evidence_quote":"The VQA baseline that shows the largest rank-correlation gain after adaptation in Table 1."},{"cited_title":"InIEEE Conference on Computer Vision and Pattern Recognition(2022)","cited_arxiv_id":null,"evidence_quote":"BasicVSR is one of the six VSR backbones in Table 2 used to demonstrate the adaptive quality loss."},{"cited_title":null,"cited_arxiv_id":null,"evidence_quote":"The transformer screen-content SR baseline that ScrVSR extends and SCISR-TTA adapts."},{"cited_title":"InEuropean conference on computer vision(2022), Springer, pp","cited_arxiv_id":null,"evidence_quote":"Supplies the fixed restoration network whose cleaned output is the reference for the non-text region loss in SCISR-TTA."},{"cited_title":"In2025 Data Compression Conference (DCC)(2025), IEEE, pp","cited_arxiv_id":null,"evidence_quote":"A competing screen-content enhancement method used as comparison in the ScrVSR and SCISR-TTA tables."}],"review_version":1}